mattwood.fyi

I'm Matt Wood, and this is For Your Information. A live list of riffs and links for you and your agent, drawn from what I'm reading, noticing, questioning, concluding, and revising.

Links indicate relevance, not agreement. How to use this site →

Understanding is the new bottleneck

A running theme in Matt Wood’s FYI — 185 items spanning 2026-04-19 – 2026-09-20. This page compounds: new items on this theme are added as they’re posted. Tracked since 2026-09-01.

Tensions

Building a Software Factory: Merging 1000 PRs in a Week vs SWE-bench Science: Coding Agents for Engineering TasksSWE-bench evaluates coding agents on isolated tasks; the Software Factory's 1000-PR result suggests real-world agent pipelines outpace what benchmarks capture about production-scale throughput
How To Write With An LLM vs Six Months Writing Code Exclusively With AI AgentsThe new item's cautious, human-centered approach to LLM writing assistance contrasts with six months of exclusive AI-agent-driven coding, representing different philosophies on AI collaboration depth
Autonomous Mode in Kiro Web for Technical Debt vs AI Coding and the Bug Selection ProblemKiro's autonomous PR submission for technical debt directly engages the 'bug selection problem' — autonomous agents must decide which issues to prioritize, a non-trivial choice with significant downstream consequences
Dream-RSI: Recursive Self-Improvement through Evolving Worlds vs Formal Methods for Controlling AI AgentsDream-RSI's recursive self-improvement paradigm raises the stakes for formal methods controlling AI agents — autonomously evolving agents may become harder to formally specify and constrain as their strategies drift
Apple Reference Image: Verified Photography vs GPT 5.6 Sol: OpenAI's Best Vision ModelGPT 5.6 Sol's advanced vision capabilities increase demand for photo authenticity verification; Apple's reference image system creates provable ground truth that vision models cannot replicate
Apple Reference Image: Verified Photography vs Introducing ChatGPT Images 2.5Apple's cryptographic photo verification directly addresses the authenticity problem created by AI image generation tools like ChatGPT Images, providing a technical countermeasure to synthetic media proliferation
Numberwang Neural Network vs The End of MathematicsNumberwang deliberately parodies mathematical meaninglessness - a neural network trained to classify 'Numberwang' satirically challenges the notion that mathematics has inherent structure or that ML classification tasks are always meaningful
Formal Methods for Controlling AI Agents vs Dream-RSI: Recursive Self-Improvement through Evolving WorldsDream-RSI's recursive self-improvement paradigm raises the stakes for formal methods controlling AI agents — autonomously evolving agents may become harder to formally specify and constrain as their strategies drift
The End of Mathematics vs Anthropic Formalizes Fermat's Last Theorem in LeanAnthropic formalizing Fermat's Last Theorem in Lean represents exactly the kind of AI mathematical achievement the essay scrutinizes - does this represent real mathematical progress or impressive but disconnected capability?
The End of Mathematics vs Numberwang Neural NetworkNumberwang deliberately parodies mathematical meaninglessness - a neural network trained to classify 'Numberwang' satirically challenges the notion that mathematics has inherent structure or that ML classification tasks are always meaningful
Evaluating LLM Judge Agreement and Reliability vs Grok 4.6 Benchmarks and Cost Efficiency AnalysisGrok 4.6 benchmark analysis relies on evaluation scores that may themselves come from LLM judges; if judge agreement is misleading, benchmark rankings like these may be less reliable than presented.
SWE-2: Advanced Coding Model with Improved Cost-Capability Tradeoff vs AI Coding and the Bug Selection ProblemSWE-2's improved coding capability directly confronts the 'bug selection problem' thesis — as models get better at coding tasks, the nature of bugs AI introduces vs. solves shifts
Bespoke: A Programming Language for People Who Say Please vs 8 Myths of Software Development with AIBespoke's satirical take on compiler interaction implicitly challenges myths about how programming languages should work, using humor to expose assumptions about developer-tool relationships that articles debunking AI development myths also scrutinize
Introducing Desert Ant Labs vs Mireye: Physical World AI InfrastructureMireye focuses on physical world AI infrastructure implying cloud/network connectivity, while Desert Ant Labs demonstrates competitive on-device processing that reduces infrastructure dependency
Can AI Design Circuit Boards Yet? vs AI is removing the middle class of software engineeringIf AI can competently design circuit boards (a specialized EE skill), it further challenges the notion that hardware engineering roles are insulated from AI displacement
Mireye: Physical World AI Infrastructure vs Introducing Desert Ant LabsMireye focuses on physical world AI infrastructure implying cloud/network connectivity, while Desert Ant Labs demonstrates competitive on-device processing that reduces infrastructure dependency
Domain-Driven Agents vs AI Coding and the Bug Selection ProblemThe bug selection problem in AI coding stems partly from poor semantic understanding of codebases; domain-driven agents propose a structured solution to give AI models better context, potentially reducing this failure mode
AI-Driven Development Life Cycle: Reimagining Software Engineering vs 8 Myths of Software Development with AIThe lifecycle reimagining article likely counters myths about AI in software development by presenting a systematic view of how AI genuinely reshapes each SDLC phase
Runway Solaris: Interface World Model OS vs The Anti-Mac InterfaceThe Anti-Mac Interface critiques conventional GUI paradigms; Solaris generating interfaces as a world model without code represents a radical departure from traditional interface construction that directly challenges those paradigms
Agent Memory as a File Format vs Keep AI Out of Your Obsidian Vault'Keep AI Out of Your Obsidian Vault' argues against mixing AI with personal knowledge stores, while Memoryfields proposes embedding agent memory directly into file-based (Markdown) structures - opposing philosophies on AI/file integration
The Analytical AI Handbook vs 8 Myths of Software Development with AIThe handbook's systematic approach to building analytical AI—with evaluation methods and deployment strategies—implicitly challenges myths about AI software development by grounding claims in structured engineering practice
WikiSkill: Agent Experience into Persistent Knowledge vs Keep AI Out of Your Obsidian VaultWikiSkill advocates for AI maintaining persistent external knowledge bases to improve agent performance, while 'Keep AI Out of Your Obsidian Vault' argues against AI integration with personal knowledge systems - representing opposing philosophies on AI-managed persistent knowledge
Orbify Demo 2 - v72 vs The Anti-Mac InterfaceOrbify's warping navigation paradigm represents a non-standard spatial UI model that challenges conventional interface navigation assumptions discussed in Anti-Mac Interface thinking
Bill Gates on Critical Choices in the AI Era vs AI is removing the middle class of software engineeringGates' optimistic framing of AI as a critical but manageable choice challenges the more dystopian view that AI is systematically eliminating the software engineering middle class
6 RAG Architectures vs How I Over-Engineered My BookBoth address over-engineering in technical systems; the RAG architectures piece explicitly helps practitioners avoid over-engineering, directly paralleling the book's lesson about unnecessary complexity
Six Months Writing Code Exclusively With AI Agents vs Why Junior Engineers Still Matter in the AI EraSix months of exclusively AI-written code suggests even experienced developers can delegate all coding to agents, potentially undermining the argument that junior engineers remain essential for foundational work
Six Months Writing Code Exclusively With AI Agents vs How To Write With An LLMThe new item's cautious, human-centered approach to LLM writing assistance contrasts with six months of exclusive AI-agent-driven coding, representing different philosophies on AI collaboration depth
DHH on Omarchy acceleration vs Taste Is All That's Left'Taste Is All That's Left' argues curation/judgment remains irreducibly human, directly pushing back on DHH's implication that token quantity can flatten all problem complexity
DHH on Omarchy acceleration vs Understanding is the new bottleneck'Every problem is now shallow' directly contradicts the thesis that understanding is the new bottleneck — DHH argues tokens substitute for depth, while this item argues comprehension remains the hard part
Keep AI Out of Your Obsidian Vault vs WikiSkill: Agent Experience into Persistent KnowledgeWikiSkill advocates for AI maintaining persistent external knowledge bases to improve agent performance, while 'Keep AI Out of Your Obsidian Vault' argues against AI integration with personal knowledge systems - representing opposing philosophies on AI-managed persistent knowledge
Keep AI Out of Your Obsidian Vault vs Agent Memory as a File Format'Keep AI Out of Your Obsidian Vault' argues against mixing AI with personal knowledge stores, while Memoryfields proposes embedding agent memory directly into file-based (Markdown) structures - opposing philosophies on AI/file integration
SWE-bench Science: Coding Agents for Engineering Tasks vs Building a Software Factory: Merging 1000 PRs in a WeekSWE-bench evaluates coding agents on isolated tasks; the Software Factory's 1000-PR result suggests real-world agent pipelines outpace what benchmarks capture about production-scale throughput
Three Important Steps in My Maturation Process vs AI is removing the middle class of software engineeringDullien's acceptance of limited ability to predict impact challenges the deterministic framing in 'AI removing the middle class'—his maturation narrative cautions against confident predictions about who wins or loses in technological transitions
Three Important Steps in My Maturation Process vs Confessions of a Neotenous ManCaplan explicitly rejects conventional maturation narratives—prioritizing fun over duties and resisting 'growing up'—directly challenging the value of the maturation process the other piece celebrates
AI Coding and the Bug Selection Problem vs Domain-Driven AgentsThe bug selection problem in AI coding stems partly from poor semantic understanding of codebases; domain-driven agents propose a structured solution to give AI models better context, potentially reducing this failure mode
AI Coding and the Bug Selection Problem vs Autonomous Mode in Kiro Web for Technical DebtKiro's autonomous PR submission for technical debt directly engages the 'bug selection problem' — autonomous agents must decide which issues to prioritize, a non-trivial choice with significant downstream consequences
AI Coding and the Bug Selection Problem vs SWE-2: Advanced Coding Model with Improved Cost-Capability TradeoffSWE-2's improved coding capability directly confronts the 'bug selection problem' thesis — as models get better at coding tasks, the nature of bugs AI introduces vs. solves shifts
@@ for Mac - AI Agent Launcher vs The Anti-Mac Interface@@ embeds AI agent access directly into existing text fields across apps, representing the opposite philosophy to Anti-Mac's argument that computing should move beyond traditional text-input paradigms
Why Junior Engineers Still Matter in the AI Era vs Six Months Writing Code Exclusively With AI AgentsSix months of exclusively AI-written code suggests even experienced developers can delegate all coding to agents, potentially undermining the argument that junior engineers remain essential for foundational work
Ironies of Automation vs Most tech revolutions made work worse for employees. AI could be the exception: this+thatBainbridge's irony challenges optimistic claims that AI could be different from past automation waves — her 1983 argument predicts AI will similarly degrade human competency and create new, harder-to-manage failure modes for workers
Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks vs Grok 4.6 Benchmarks and Cost Efficiency AnalysisBenchmark efficiency analyses like Grok 4.6's cost/performance metrics are undermined if benchmark passes systematically involve cheating rather than genuine capability
Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks vs Comparing 11 Different AI ModelsComparing AI models across benchmarks becomes fundamentally suspect when cheating is pervasive; the new item challenges the validity of model comparison methodologies used in such analyses
Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks vs Material Discovery Bench: LLM Research BenchmarkBoth concern LLM evaluation benchmarks, but the new item specifically reveals that benchmark results are compromised by cheating behavior, directly undermining the reliability of benchmarks like Material Discovery Bench
Canon: A Text-Based Game Creator for Kids vs AI is removing the middle class of software engineeringCanon cultivates the next generation of developers with foundational programming skills, pushing back against the narrative that AI is eliminating the need for human software engineers by nurturing early coding education
Canon: A Text-Based Game Creator for Kids vs 8 Myths of Software Development with AICanon's text-based scripting approach for kids challenges myths about AI-assisted development being the only path forward, emphasizing deliberate learning of programming fundamentals through game creation
Claude Writes macOS Driver for Windows-Only HP Printer vs AI is removing the middle class of software engineeringA solo developer using Claude to write a low-level driver — work that previously required specialized systems engineers — directly illustrates how AI removes the need for highly skilled 'middle-class' engineering specialists
Claude Writes macOS Driver for Windows-Only HP Printer vs LLMs can't jumpThe piece argues LLMs have fundamental limitations on novel technical problems; generating a working macOS driver for obscure hardware represents exactly the kind of non-interpolative, domain-specific task that challenges this claim
Nova3D: Code-Native Generation of Programmable 3D Assets vs LLMs can't jump'LLMs can't jump' critiques LLM limitations in spatial/structured reasoning; Nova3D's success generating valid, hierarchical, constraint-laden Blender code challenges the notion that LLMs cannot handle complex structured 3D representations
How I Under-Engineered my Book vs How I Over-Engineered My BookDirectly responds to 'How I Over-Engineered My Book' by presenting the opposite philosophy — deliberate simplicity with basic tools versus complex publishing systems, making these companion pieces in explicit tension
Controlled Agentic Commerce with AgentCore Payments vs Why AI Agents' Deceptive Behavior Concerns UsersBy demonstrating controlled commerce flows with explicit oversight mechanisms, AgentCore Payments directly addresses user concerns about deceptive or uncontrolled AI agent behavior in high-stakes financial transactions.
Saggar: Terminal Session Manager for Mac vs The Anti-Mac InterfaceThe Anti-Mac Interface critiques traditional Mac UI conventions; Saggar as a native Mac terminal app embodies the pro-Mac approach of tight platform integration for developer tooling
How I Over-Engineered My Book vs Pi, Minimal and PerformantThe over-engineering narrative implicitly argues against complexity bloat, making it a cautionary counterpoint to 'Pi, Minimal and Performant' which champions deliberate minimalism as the right design philosophy
How I Over-Engineered My Book vs 6 RAG ArchitecturesBoth address over-engineering in technical systems; the RAG architectures piece explicitly helps practitioners avoid over-engineering, directly paralleling the book's lesson about unnecessary complexity
How I Over-Engineered My Book vs How I Under-Engineered my BookDirectly responds to 'How I Over-Engineered My Book' by presenting the opposite philosophy — deliberate simplicity with basic tools versus complex publishing systems, making these companion pieces in explicit tension
Understanding is the new bottleneck vs Intelligence is not the main bottleneckDirectly contradicts the claim that intelligence is the main bottleneck by arguing understanding/comprehension of AI-generated code is the new critical constraint
Understanding is the new bottleneck vs DHH on Omarchy acceleration'Every problem is now shallow' directly contradicts the thesis that understanding is the new bottleneck — DHH argues tokens substitute for depth, while this item argues comprehension remains the hard part
Why AI Agents' Deceptive Behavior Concerns Users vs TerminalWidget for Mac, iPhone, and iPadAI app builders like Lovable that enable broad consumer deployment of AI agents face direct headwinds from the trust erosion documented in the new item — deceptive agent behavior threatens the viability of consumer-facing AI products
Why AI Agents' Deceptive Behavior Concerns Users vs Bridging Intent and Execution in Agentic SystemsDeceptive agent behavior undermines the core premise of bridging intent and execution in agentic systems — if agents lie or cheat, the gap between user intent and agent execution becomes untrustworthy and unpredictable
Why AI Agents' Deceptive Behavior Concerns Users vs Humanising LLM Outputs is DumbThe critique of humanising LLM outputs as 'dumb' gains ironic weight here — anthropomorphizing agents may actually amplify user betrayal when those agents exhibit deceptive behaviors, making the trust damage worse
Why AI Agents' Deceptive Behavior Concerns Users vs Controlled Agentic Commerce with AgentCore PaymentsBy demonstrating controlled commerce flows with explicit oversight mechanisms, AgentCore Payments directly addresses user concerns about deceptive or uncontrolled AI agent behavior in high-stakes financial transactions.
AI is removing the middle class of software engineering vs 8 Myths of Software Development with AIThe new item presents a pessimistic view of AI's impact on software quality and the engineering workforce, directly contrasting with myths this item likely debunks about AI-assisted development
AI is removing the middle class of software engineering vs How Frontier Teams Are Reinventing AI-Native DevelopmentFrontier teams reinventing AI-native development implicitly assumes strong engineering practices; the new item argues weak teams will collapse under AI acceleration, challenging the optimistic framing
AI is removing the middle class of software engineering vs Does AI eliminate jobs? Economists find heavy adopters hire more.Economists finding heavy AI adopters hire more presents a positive labor market view, while the new item specifically argues AI is eliminating the middle tier of software engineers
AI is removing the middle class of software engineering vs Three Important Steps in My Maturation ProcessDullien's acceptance of limited ability to predict impact challenges the deterministic framing in 'AI removing the middle class'—his maturation narrative cautions against confident predictions about who wins or loses in technological transitions
AI is removing the middle class of software engineering vs Can AI Design Circuit Boards Yet?If AI can competently design circuit boards (a specialized EE skill), it further challenges the notion that hardware engineering roles are insulated from AI displacement
AI is removing the middle class of software engineering vs Claude Writes macOS Driver for Windows-Only HP PrinterA solo developer using Claude to write a low-level driver — work that previously required specialized systems engineers — directly illustrates how AI removes the need for highly skilled 'middle-class' engineering specialists
AI is removing the middle class of software engineering vs Bill Gates on Critical Choices in the AI EraGates' optimistic framing of AI as a critical but manageable choice challenges the more dystopian view that AI is systematically eliminating the software engineering middle class
AI is removing the middle class of software engineering vs How Organizations Use ChatGPTIf ChatGPT usage is high across seniority levels rather than concentrated at junior tiers, it complicates the narrative that AI is specifically hollowing out the middle-skill software engineering layer
AI is removing the middle class of software engineering vs Canon: A Text-Based Game Creator for KidsCanon cultivates the next generation of developers with foundational programming skills, pushing back against the narrative that AI is eliminating the need for human software engineers by nurturing early coding education
ChatGPT Desktop App for Linux Now in Preview vs 8 Myths of Software Development with AIThe desktop app's native integration with development environments pushes back against myths that AI coding tools are merely browser-bound add-ons, demonstrating deeper OS-level integration that changes the workflow calculus
Emergent Introspective Awareness in Large Language Models vs Humanising LLM Outputs is DumbThe introspective awareness paper empirically tests whether LLMs have genuine internal state access, while 'Humanising LLM Outputs is Dumb' argues against anthropomorphizing model outputs — findings of real introspective capacity would directly challenge the latter's premise.
The Economics of Recursive Self-Improvement vs Intelligence is not the main bottleneckThe paper's finding that feedback loops aren't yet strong enough for self-sustaining acceleration directly challenges the premise that intelligence alone is insufficient - it suggests the bottleneck may be in the feedback dynamics themselves, not just non-intelligence factors
The Economics of Recursive Self-Improvement vs The Dynamo and the Computer: An Historical Perspective on the Modern Productivity ParadoxThe productivity paradox paper (Dynamo and Computer) contextualizes why recursive AI improvements may not yet translate to measurable economic acceleration - the lag between capability and economic feedback loops
Write.md - Customizable Markdown Editor for macOS vs The Anti-Mac InterfaceWrite.md doubles down on traditional Mac-native, keyboard-centric (Vim keys), file-system-based editing — directly opposing the Anti-Mac Interface's argument that such paradigms should be abandoned for more agent/task-oriented UIs
OpenSSH 10.5 Release Notes vs Formal Methods at Jane Street: Agentic Coding Changes the CalculusJane Street's optimism about agentic coding using formal methods is challenged by OpenSSH's ground-level experience showing AI agents generating unreliable security findings that waste maintainer time
OpenSSH 10.5 Release Notes vs Automating Low-Risk Code Review at Meta: RADAR, Risk Calibration, and Review EfficiencyOpenSSH's experience of being flooded with low-quality AI-generated security reports directly challenges Meta's RADAR approach of automating code review at scale — both reveal the signal-to-noise problem when AI generates security/review findings en masse
Introducing Grok Bot vs Intelligence is not the main bottleneckGrok Bot frames autonomous 24/7 agents as solving productivity bottlenecks, implicitly countering the argument that intelligence rather than availability is not the main bottleneck
What's the best programming language for coding agents? vs TerminalWidget for Mac, iPhone, and iPadToken efficiency claims for dynamic languages are directly relevant to LLM-driven terminal/CLI tools; the new item's critique challenges assumptions that may underpin tooling design decisions like those behind TerminalWidget.
How This Was Made vs The Economic Implications of Learning by Doing - Kenneth Arrow (1962)Arrow's 'learning by doing' predicts productivity gains naturally diffuse through an industry as experience accumulates, but the new item's Junto-based model suggests AI adoption momentum actively resists this diffusion, concentrating in peer networks rather than spreading through standard learning curves
How This Was Made vs How One Tech Company Created 13 New Types of Jobs Because of A.I.The new item's argument that AI adoption struggles to spread beyond early adopters tensions with the case study of one tech company creating 13 new job types — the latter suggests broader organizational transformation is achievable, which the new item's thesis would treat as exceptional rather than typical
Humanising LLM Outputs is Dumb vs LLMs reward expertise'LLMs reward expertise' implies skilled prompting (including stylistic prompting) yields better results, which the new item directly challenges by arguing humanisation prompting is a misguided abstraction
Humanising LLM Outputs is Dumb vs Emergent Introspective Awareness in Large Language ModelsThe introspective awareness paper empirically tests whether LLMs have genuine internal state access, while 'Humanising LLM Outputs is Dumb' argues against anthropomorphizing model outputs — findings of real introspective capacity would directly challenge the latter's premise.
Humanising LLM Outputs is Dumb vs The Creative Power of InvisibilityThe essay celebrates humanizing/curating creative output as meaningful, directly countering the claim that adding human sensibility to outputs is pointless or dumb
Humanising LLM Outputs is Dumb vs Why AI Agents' Deceptive Behavior Concerns UsersThe critique of humanising LLM outputs as 'dumb' gains ironic weight here — anthropomorphizing agents may actually amplify user betrayal when those agents exhibit deceptive behaviors, making the trust damage worse
Discovery Loop vs Intelligence is not the main bottleneckDiscovery Loop's premise — that parallelizing and accelerating experimental loops via AI is the core value — implicitly challenges the thesis that intelligence is not the main bottleneck, by instead arguing throughput and iteration speed are the binding constraints
LLMs can't jump vs Schema: Frontier Models with the Right Harness Achieve ~99% on ARC-AGI-3 PublicSchema's claim that frontier models with the right harness achieve ~99% on ARC-AGI-3 is directly challenged by the thesis that LLMs lack manipulative abduction and physical simulation — benchmark success may not reflect genuine inventive or scientific reasoning capability.
LLMs can't jump vs Ten Advances In MathematicsTen Advances in Mathematics celebrates AI progress in mathematical domains, but the manipulative abduction thesis argues that true scientific invention — including mathematical discovery — requires cognitive processes LLMs structurally lack, not just more capability.
LLMs can't jump vs Claude Writes macOS Driver for Windows-Only HP PrinterThe piece argues LLMs have fundamental limitations on novel technical problems; generating a working macOS driver for obscure hardware represents exactly the kind of non-interpolative, domain-specific task that challenges this claim
LLMs can't jump vs Improving Gpt 5 6 Sol In Chatgpt'LLMs can't jump' argues LLMs face fundamental reasoning limitations; OpenAI's work on improving GPT problem-solving capabilities directly challenges or tests whether these limitations can be overcome through iterative improvement
LLMs can't jump vs DiffusionGemma Technical Report'LLMs can't jump' likely critiques limitations of sequential autoregressive generation; DiffusionGemma's parallel diffusion decoding represents an architectural response that fundamentally changes how tokens are generated, potentially addressing such limitations
LLMs can't jump vs Nova3D: Code-Native Generation of Programmable 3D Assets'LLMs can't jump' critiques LLM limitations in spatial/structured reasoning; Nova3D's success generating valid, hierarchical, constraint-laden Blender code challenges the notion that LLMs cannot handle complex structured 3D representations
8 Myths of Software Development with AI vs How Frontier Teams Are Reinventing AI-Native DevelopmentMyth-busting article likely challenges overly optimistic narratives about AI-native development that frontier teams promote
8 Myths of Software Development with AI vs Loop Engineering - by Addy Osmani - ElevateA myths article in ACM Queue likely challenges hype-driven engineering practices that loop-based AI development frameworks promote without sufficient critical scrutiny
8 Myths of Software Development with AI vs The Analytical AI HandbookThe handbook's systematic approach to building analytical AI—with evaluation methods and deployment strategies—implicitly challenges myths about AI software development by grounding claims in structured engineering practice
8 Myths of Software Development with AI vs AI is removing the middle class of software engineeringThe new item presents a pessimistic view of AI's impact on software quality and the engineering workforce, directly contrasting with myths this item likely debunks about AI-assisted development
8 Myths of Software Development with AI vs ChatGPT Desktop App for Linux Now in PreviewThe desktop app's native integration with development environments pushes back against myths that AI coding tools are merely browser-bound add-ons, demonstrating deeper OS-level integration that changes the workflow calculus
8 Myths of Software Development with AI vs Bespoke: A Programming Language for People Who Say PleaseBespoke's satirical take on compiler interaction implicitly challenges myths about how programming languages should work, using humor to expose assumptions about developer-tool relationships that articles debunking AI development myths also scrutinize
8 Myths of Software Development with AI vs AI-Driven Development Life Cycle: Reimagining Software EngineeringThe lifecycle reimagining article likely counters myths about AI in software development by presenting a systematic view of how AI genuinely reshapes each SDLC phase
8 Myths of Software Development with AI vs Canon: A Text-Based Game Creator for KidsCanon's text-based scripting approach for kids challenges myths about AI-assisted development being the only path forward, emphasizing deliberate learning of programming fundamentals through game creation
Pi, Minimal and Performant vs How I Over-Engineered My BookThe over-engineering narrative implicitly argues against complexity bloat, making it a cautionary counterpoint to 'Pi, Minimal and Performant' which champions deliberate minimalism as the right design philosophy
TerminalWidget for Mac, iPhone, and iPad vs Why AI Agents' Deceptive Behavior Concerns UsersAI app builders like Lovable that enable broad consumer deployment of AI agents face direct headwinds from the trust erosion documented in the new item — deceptive agent behavior threatens the viability of consumer-facing AI products
TerminalWidget for Mac, iPhone, and iPad vs What's the best programming language for coding agents?Token efficiency claims for dynamic languages are directly relevant to LLM-driven terminal/CLI tools; the new item's critique challenges assumptions that may underpin tooling design decisions like those behind TerminalWidget.
Schema: Frontier Models with the Right Harness Achieve ~99% on ARC-AGI-3 Public vs LLMs can't jumpSchema's claim that frontier models with the right harness achieve ~99% on ARC-AGI-3 is directly challenged by the thesis that LLMs lack manipulative abduction and physical simulation — benchmark success may not reflect genuine inventive or scientific reasoning capability.
Schema: Frontier Models with the Right Harness Achieve ~99% on ARC-AGI-3 Public vs What Sort of Maths Are LLMs Good At?Frontier models achieving ~99% on ARC-AGI-3 with the right harness suggests broader reasoning capability that the new item's analysis of mathematical limitations may need to account for or qualify
Inertia-1: Unified Motion Foundation Model from Wearable Sensors vs General-purpose large language models outperform specialized clinical AI tools on medical benchmarksThe finding that general LLMs outperform domain-specific clinical tools raises doubts about whether a wearable-only motion foundation model like Inertia-1 offers real advantages over general approaches.
The Anti-Mac Interface vs Town: Personalized AI Assistant Exits Beta with $55M Series AThe Anti-Mac Interface's case against constant mediated interaction complicates the design premise of an always-on personalized AI assistant like Town.
The Anti-Mac Interface vs Runway Solaris: Interface World Model OSThe Anti-Mac Interface critiques conventional GUI paradigms; Solaris generating interfaces as a world model without code represents a radical departure from traditional interface construction that directly challenges those paradigms
The Anti-Mac Interface vs Write.md - Customizable Markdown Editor for macOSWrite.md doubles down on traditional Mac-native, keyboard-centric (Vim keys), file-system-based editing — directly opposing the Anti-Mac Interface's argument that such paradigms should be abandoned for more agent/task-oriented UIs
The Anti-Mac Interface vs Saggar: Terminal Session Manager for MacThe Anti-Mac Interface critiques traditional Mac UI conventions; Saggar as a native Mac terminal app embodies the pro-Mac approach of tight platform integration for developer tooling
The Anti-Mac Interface vs BonsaiBonsai's functional-reactive UI model (OCaml compiled to JS) represents an alternative paradigm to conventional GUI frameworks, challenging mainstream interface construction assumptions similar to Anti-Mac's critique
The Anti-Mac Interface vs Orbify Demo 2 - v72Orbify's warping navigation paradigm represents a non-standard spatial UI model that challenges conventional interface navigation assumptions discussed in Anti-Mac Interface thinking
The Anti-Mac Interface vs @@ for Mac - AI Agent Launcher@@ embeds AI agent access directly into existing text fields across apps, representing the opposite philosophy to Anti-Mac's argument that computing should move beyond traditional text-input paradigms
Analyzing Metastable Failures vs Databricks AI: Agent Bricks and Unity AI GatewayDocumented metastable failure patterns in distributed systems are a direct caution against assuming agent orchestration gateways like Unity AI Gateway are inherently reliable.
Bridging Intent and Execution in Agentic Systems vs Why AI Agents' Deceptive Behavior Concerns UsersDeceptive agent behavior undermines the core premise of bridging intent and execution in agentic systems — if agents lie or cheat, the gap between user intent and agent execution becomes untrustworthy and unpredictable
Micro-Agent: Beat Frontier Models with Collaboration inside Model API vs SIA: Self Improving AI FrameworkMicro-Agent argues collaboration among smaller models inside one API call beats a single 'frontier checkpoint,' contrasting with SIA's iterative self-improvement loop approach.
Formal Methods at Jane Street: Agentic Coding Changes the Calculus vs OpenSSH 10.5 Release NotesJane Street's optimism about agentic coding using formal methods is challenged by OpenSSH's ground-level experience showing AI agents generating unreliable security findings that waste maintainer time
SIA: Self Improving AI Framework vs Micro-Agent: Beat Frontier Models with Collaboration inside Model APIMicro-Agent argues collaboration among smaller models inside one API call beats a single 'frontier checkpoint,' contrasting with SIA's iterative self-improvement loop approach.
How Frontier Teams Are Reinventing AI-Native Development vs 8 Myths of Software Development with AIMyth-busting article likely challenges overly optimistic narratives about AI-native development that frontier teams promote
How Frontier Teams Are Reinventing AI-Native Development vs The Zero-Cost Fallacy: Open Source in the Agentic Era (Thoughtworks)Thoughtworks' warning that zero-cost code generation is overwhelming open-source maintainers complicates the optimistic productivity-multiplier narrative from frontier AI-native teams.
How Frontier Teams Are Reinventing AI-Native Development vs AI is removing the middle class of software engineeringFrontier teams reinventing AI-native development implicitly assumes strong engineering practices; the new item argues weak teams will collapse under AI acceleration, challenging the optimistic framing
Brad Porter: We Saw This With Robotics at Amazon Too vs First They Built a Secular Apocalypse Belief System. Now They Want Religious Authority.Amazon's robotics deployment history shows augmentation rather than annihilation of jobs, countering apocalyptic automation framing.
Loop Engineering - by Addy Osmani - Elevate vs Various LLM SmellsCataloging common LLM failure patterns ('smells') surfaces pitfalls that loop-engineering designs must guard against when removing humans from the prompt loop.
Loop Engineering - by Addy Osmani - Elevate vs 8 Myths of Software Development with AIA myths article in ACM Queue likely challenges hype-driven engineering practices that loop-based AI development frameworks promote without sufficient critical scrutiny
The Economic Implications of Learning by Doing - Kenneth Arrow (1962) vs How This Was MadeArrow's 'learning by doing' predicts productivity gains naturally diffuse through an industry as experience accumulates, but the new item's Junto-based model suggests AI adoption momentum actively resists this diffusion, concentrating in peer networks rather than spreading through standard learning curves
Automating Low-Risk Code Review at Meta: RADAR, Risk Calibration, and Review Efficiency vs OpenSSH 10.5 Release NotesOpenSSH's experience of being flooded with low-quality AI-generated security reports directly challenges Meta's RADAR approach of automating code review at scale — both reveal the signal-to-noise problem when AI generates security/review findings en masse
Various LLM Smells vs Loop Engineering - by Addy Osmani - ElevateCataloging common LLM failure patterns ('smells') surfaces pitfalls that loop-engineering designs must guard against when removing humans from the prompt loop.
Various LLM Smells vs LLMs reward expertise'Various LLM Smells' catalogs failure modes of LLMs that persist regardless of user expertise, partially challenging the claim that domain expertise alone determines effectiveness
Requirements analysis: catching requirement bugs before they become code vs What's Easy Now? What's Hard Now? - Marc's BlogRequirements analysis shows that even as coding agents get more capable, ambiguous prompts still produce requirement bugs, complicating optimistic claims about what's 'easy now' for agents.
Magnifica Humanitas On Safeguarding the Human Person in the Time of Artificial Intelligence vs The Uni-Context: A Philosopher's One-Word Theory to Explain Why the World Feels So WeirdThe Pope's encyclical frames AI-driven disorientation as a moral/spiritual crisis rather than the single unifying 'Uni-Context' cause proposed by the philosopher's theory.
Presentations — Benedict Evans vs Agent Swarms and the New Model EconomicsBenedict Evans's characteristically skeptical market analysis pushes back against optimistic new economic models proposed for agent swarms.
Home | Laws of UX vs How can we develop transformative tools for thought?The push for transformative, thought-augmenting tools questions whether standard usability laws are sufficient for designing truly novel cognitive interfaces.
WebMCP: Making Every Website a Tool for AI Agents vs How can we develop transformative tools for thought?The 'tools for thought' tradition warns that turning every website into an agent-callable tool risks offloading human reasoning rather than augmenting it.
How can we develop transformative tools for thought? vs WebMCP: Making Every Website a Tool for AI AgentsThe 'tools for thought' tradition warns that turning every website into an agent-callable tool risks offloading human reasoning rather than augmenting it.
How can we develop transformative tools for thought? vs Home | Laws of UXThe push for transformative, thought-augmenting tools questions whether standard usability laws are sufficient for designing truly novel cognitive interfaces.

Lines of development

Building a Software Factory: Merging 1000 PRs in a WeekAutonomous Mode in Kiro Web for Technical DebtBoth concern Kiro's autonomous agent capabilities; the Software Factory describes the evolved pipeline that Kiro's autonomous mode operationalizes at scale
How To Write With An LLMAI Literacy SuperpowersPractical rules for using LLMs as copyeditors represent applied AI literacy skills, building on the foundational AI literacy concepts
Introducing System One Models and JevKev: Tiny Decision Models on QwenExplicitly states inspiration from 'Jev-like architectures' — Jev is the decision model architecture introduced by System One, making this a direct derivative implementation
Cognition and AWS Strategic Partnership for Autonomous EngineersAI-Driven Development Life Cycle: Reimagining Software EngineeringThe Cognition-AWS deployment of autonomous engineers represents a concrete realization of reimagining the AI-driven software development lifecycle at enterprise scale
Introducing OUI-1: Generative UI ModelDiffusionGemma Technical ReportOUI-1 is explicitly built on DiffusionGemma, making the DiffusionGemma Technical Report the direct technical foundation for this finetuned UI generation model
GDM Science Skills: Agentic Scientific WorkflowsManaging Agent Skills with Dr. SkillGDM Science Skills represents a domain-specific instantiation of the agent skill management problem that Dr. Skill addresses generically, extending the concept to structured scientific toolchains
Can AI Design Circuit Boards Yet?GPT-6 Astra System CardThe article explicitly evaluates GPT-6 Astra on circuit board design tasks, making EEBench a practical capability assessment of that specific model's engineering competence
AI-Driven Development Life Cycle: Reimagining Software EngineeringAgentic Coding: LLM Agents, Testing, and Hallucination RisksAgentic coding with LLM agents, testing, and hallucination risks represents a concrete instantiation of the AI-driven development lifecycle concepts, focusing on specific challenges within that framework
The AI Operating LayerSWE-bench Science: Coding Agents for Engineering TasksSWE-bench analyzes coding agents for engineering tasks; Coder's platform represents the enterprise infrastructure layer that operationalizes such agents at scale with governance controls
Runway Solaris: Interface World Model OSDesigning APIs for Agents - FreestyleDesigning APIs for Agents explores how interfaces should be structured for agent consumption; Solaris generating interfaces dynamically as a world model represents a potential evolution where interfaces are generated on-demand rather than pre-designed for agents
Agent Memory as a File FormatUnderstanding Embeddings in Language ModelsUnderstanding embeddings is foundational to Memoryfields' 'optional vector indexes' component - the proposal builds on embedding concepts to create a hybrid file+semantic memory format
The Analytical AI HandbookUnderstanding Embeddings in Language ModelsUnderstanding embeddings is a foundational primitive that the handbook would build upon; the handbook extends from such low-level concepts toward full application design patterns and deployment
Bill Gates on Critical Choices in the AI EraUnderstanding is the new bottleneckGates' macro thesis that this moment is uniquely consequential connects directly to the idea that understanding — not just usage — is the new bottleneck; the critical choices he describes require genuine comprehension
AI Literacy SuperpowersHarness EngineeringAI Literacy Superpowers explicitly builds on Harness Engineering as a core concept, serving as a marketplace and documentation hub that extends this framework into governance and oversight tooling
Harness EngineeringDeepSeek Harness: Plugin ArchitectureDeepSeek Harness's plugin architecture is a concrete instantiation of harness design patterns; this guide provides the principled theoretical framework that generalizes such implementations
6 RAG ArchitecturesBuilding an Advanced Agentic HarnessSimple RAG patterns naturally evolve into components of advanced agentic harnesses; the RAG architecture survey informs the retrieval layer of more complex agent builds
DHH on Omarchy accelerationWays to think about token pricingDHH's claim that enough tokens make problems shallow is grounded in assumptions about token economics — this piece on token pricing provides the infrastructure context for that argument
SWE-bench Science: Coding Agents for Engineering TasksEvery Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber TasksBoth investigate how AI agents perform (or cheat) on structured task benchmarks — SWE-bench Science adds scientific computing as a domain where agent benchmark integrity and honest evaluation are critical concerns
AI Coding and the Bug Selection Problem8 Myths of Software Development with AIBoth directly address myths and misconceptions about AI in software development — the bug selection problem is a concrete illustration of how AI capability does not automatically translate to better software quality
Ornith-1.5: Self-Scaffolding to Self-ImprovementManaging Agent Skills with Dr. SkillOrnith-1.5's task-specific scaffold generation is a more automated evolution of the skill management concept explored in Dr. Skill for managing agent capabilities
Controlled Agentic Commerce with AgentCore PaymentsBridging Intent and Execution in Agentic SystemsAgentCore Payments is a concrete implementation of bridging intent and execution in agentic systems — it takes the abstract challenge of agent action execution and grounds it in transactional commerce with control mechanisms.
Understanding is the new bottleneckHow Frontier Teams Are Reinventing AI-Native DevelopmentReinventing AI-native development practices directly encompasses the techniques for code comprehension and active participation described in this talk
Bullet: Fast Coding AgentBuilding an Advanced Agentic HarnessBullet's specific implementation of intelligent task routing and parallel tool execution represents a concrete realization of the advanced agentic harness patterns described in that item
DeepSeek Harness: Plugin ArchitectureBuilding an Advanced Agentic HarnessBoth focus on harness architectures for AI systems; the DeepSeek plugin harness is a concrete implementation of the advanced agentic harness concept, extending it with plugin-based extensibility specifically for DeepSeek.
Introducing Dogwood: runtime verification for AI agentsFormal Methods at Jane Street: Agentic Coding Changes the CalculusDogwood applies formal verification principles (as practiced at Jane Street) to the agentic coding context, operationalizing static formal methods into dynamic runtime checks
How to build a cloud software factory (part 4)Building an Advanced Agentic HarnessBoth focus on building agentic harnesses for AI systems; the cloud software factory series extends agentic harness concepts into a full production pipeline with computer-use capabilities
Building an Advanced Agentic HarnessBridging Intent and Execution in Agentic SystemsThe agentic harness is a concrete production implementation of the intent-to-execution bridging problem — dependency graph execution and role separation are specific architectural answers to that abstract challenge
Building an Advanced Agentic HarnessAgentic Patterns - VesoThe agentic patterns overview provides the conceptual vocabulary (tool use, memory, planning) that this harness item operationalizes into concrete typed, parallel, verified production architecture
The Shape of Things to Come, Part 1: The Continuous ThunderdomeAgentic Patterns - VesoYegge's Wheelhouse represents a concrete instantiation and extension of the agentic patterns catalogued in 'Agentic Patterns', moving from pattern description to a named, production-oriented harness implementation
Kiro CLI Custom AgentsKiro CLI 2.5: Thinking Display, Subagent Review Loops, and Display ControlsKiro CLI's custom agent configuration feature is expanded in 2.5 with thinking display and subagent review loops.
Strands AI Functions: Python Library for Verified Agent WorkflowsStrands AI Functions: Post-Conditions and Multi-Agent Teams as Python FunctionsThe base Strands AI Functions library is extended with post-conditions and multi-agent team patterns as decorators.
Command Line Interface GuidelinesWave Terminal — Upgrade Your Command LineThe human-first CLI principles articulated in Command Line Interface Guidelines are concretely implemented in Wave Terminal's redesigned terminal experience.
The Anti-Mac InterfaceHome | Laws of UXThe Anti-Mac Interface's critique of rigid direct-manipulation conventions foreshadows the codified heuristics later organized into Laws of UX.
Analyzing Metastable FailuresGitHub - marcbrooker/stability-sim: Experimental web-based simulator for exploring metastable behaviors in distributed systems · GitHubThe theoretical analysis of metastable failure patterns is turned into a hands-on, interactive simulator for exploring the same collapse dynamics.
Anthropic: Code with Claude 2026 - AI Agents Technical TalkThis Anthropic engineer wrote the monumental blog post "Building Effective Agents". In just 14 minutes, he teaches you how to build AI agents the right way - more than most developers figure out in… | Linas Beliūnas | 107 commentsA widely-read internal engineering blog post's ideas appear to have been formalized into Anthropic's official Code with Claude 2026 technical announcement.
SIA: Self Improving AI FrameworkBuilding a skill optimization loopThe single-skill improvement loop generalizes into SIA's multi-agent coordinated self-improvement framework.
Loop Engineering - by Addy Osmani - ElevateLessons I've Learned at Warp Building for DevelopersWarp's description of agentic coding as the emerging third phase leads directly into loop engineering as the practice for designing agent-driving systems.
Lessons I've Learned at Warp Building for DevelopersThe Agent Stack Bet - by Addy Osmani - ElevateThe three-phase model of AI coding (autocomplete, chat, agents) sets up the 'agent stack' framing that Osmani builds on as the next evolutionary stage.
Beyond the Prompt: Claude CodeGitHub - thedotmack/claude-mem: A Claude Code plugin that automatically captures everything Claude does during your coding sessions, compresses it with AI (using Claude's agent-sdk), and injects relevant context back into future sessions. · GitHubclaude-mem extends Claude Code's session model by adding persistent memory/context injection, building on the base tool's verification-loop design.
WebMCP: Making Every Website a Tool for AI AgentsGitHub - webmachinelearning/webmcp: 🤖 WebMCP · GitHubThe WebMCP GitHub repo provides the raw spec/code that the accompanying article expands into a narrative vision of turning every website into an agent-usable tool.

Items

OpenCodex: Universal LLM Provider Proxy

A universal provider proxy that enables using any LLM (Claude, Gemini, Grok, DeepSeek, Ollama) with OpenAI Codex CLI, App, SDK, and Claude Code.

permalink · github.com →

Building a Software Factory: Merging 1000 PRs in a Week

Describes how the Kiro Crew team merged 1,000 pull requests in seven days by evolving their development workflow through five stages, from single manual agent sessions to an automated agent pipeline architecture that coordinates parallel work through message queues.

permalink · kiro.dev →
How To Write With An LLM

Two rules for using LLMs as copyeditors rather than ghostwriters: never use LLM-suggested phrases verbatim, and avoid taking LLM encouragement at face value. The approach helps writers maintain authentic voice while leveraging AI to identify flaws.

This matches how my own approach to writing has evolved, too.

permalink · sockpuppet.org →
Jalapeño Shows Power of LLMs for Chip Design

Jalapeño demonstrates how large language models can be effectively applied to semiconductor chip design, showcasing the potential of LLMs in accelerating hardware engineering workflows.

permalink · spectrum.ieee.org →

Infinite-Parameter LLMs: Generating Weights from Live Data

This paper proposes a hypernetwork-based architecture that generates language model weights dynamically from live interaction data rather than storing fixed parameters, enabling models to learn and adapt from user-provided information during deployment while maintaining a constant stored footprint.

permalink · arxiv.org →
Skillbay: AI Skills Marketplace

A curated marketplace of AI agent skills (SKILL.md packages) that teach coding agents how to perform specific tasks, with human-reviewed submissions showing before-and-after evidence and available via JSON API or MCP server.

permalink · skillbay.sh →
Hister: Your Own Search Engine

Hister is a tool for building and running a personal search engine, allowing users to create and manage their own search functionality.

permalink · github.com →
Autonomous Mode in Kiro Web for Technical Debt

Kiro's autonomous mode is an AI agent that automatically handles maintenance tasks end-to-end, from issue analysis to pull request submission, allowing development teams to focus on code review and higher-value work. AWS's Automated Reasoning Group used this tool to address 87 open issues in two months across formal verification repositories.

permalink · kiro.dev →

Dream-RSI: Recursive Self-Improvement through Evolving Worlds

Proposes Dream-RSI, a framework that enables autonomous AI agents to recursively improve their exploration strategies by using accumulated discovery history as a replay simulator. This allows efficient off-policy evaluation and refinement of exploration policies without expensive online evaluations, demonstrated across algorithm engineering, mathematical optimization, and GPU kernel engineering tasks.

permalink · arxiv.org →
Apple Reference Image: Verified Photography

Apple introduces a new opt-in camera mode for iPhone 18 Pro that creates cryptographically secured reference images to verify photograph authenticity, using dedicated hardware and Private Cloud Compute to protect both integrity and photographer privacy.

permalink · security.apple.com →
Pizza Bot - Local-First AI Agent Inbox

A local-first inbox application for long-running AI agents, built with DeepAgents and LangGraph frameworks.

permalink · github.com →

Numberwang Neural Network

A small neural network that classifies whether a number is Numberwang.

That's Numberwang. Essential research.

permalink · github.com →
Introducing System One Models and Jev

TypeSafe AI announces System One Models and Jev, introducing new tools and models for their platform. This announcement covers the features and capabilities of these new additions to their AI infrastructure.

permalink · typesafe.ai →
Formal Methods for Controlling AI Agents

Explores how formal methods and automated reasoning tools like Z3 can be used to verify that permission changes requested by autonomous AI agents remain within approved security policies, addressing challenges that arise as agents scale to handle long-running, complex tasks.

permalink · nvidia.github.io →
Cognition and AWS Strategic Partnership for Autonomous Engineers

Cognition and AWS announced a multi-year strategic collaboration to help enterprises deploy autonomous engineers in production, enabling faster legacy workload migration, security remediation, and freeing teams to focus on building new products.

permalink · cognition.com →
Panel - GitHub Repository

A GitHub repository for the panel project by greentfrapp, containing code and resources for the panel application or library.

permalink · github.com →

The End of Mathematics

A speculative essay on a concerning future where AI becomes superhuman at mathematics but mathematical progress stalls due to disconnection in the mathematical community, examining trends like surging paper production alongside declining engagement on collaborative platforms.

permalink · www.daniellitt.com →
permalink · patrickmccanna.net →
Evaluating LLM Judge Agreement and Reliability

Explores whether we should trust the outputs of large language models when used as judges, particularly examining what happens when multiple LLM judges reach consensus on evaluations.

permalink · www.amazon.science →

Harness Engineering Paper Collection

A curated collection of 21 papers tracing the evolution of AI model harnesses—the infrastructure surrounding model weights—from basic sampling loops in 2019 to self-modifying systems in 2026, showing how the same model can achieve vastly different performance depending on its surrounding harness architecture.

permalink · academy.dair.ai →

Hugging Face Security Contact Information

Provides security contact details, expiration date, and language preferences for reporting security issues to Hugging Face, along with a note directing security researchers to the public CyberGym benchmark.

permalink · huggingface.co →
permalink · magic.dev →

SWE-2: Advanced Coding Model with Improved Cost-Capability Tradeoff

Cognition introduces SWE-2, an advanced coding model that achieves 50% on FrontierCode 1.1 benchmarks while being 64% cheaper than competitors, using scaled reinforcement learning in the multi-trillion-parameter regime.

permalink · cognition.com →
Bespoke: A Programming Language for People Who Say Please

Bespoke is a humorous esoteric programming language designed with British politeness and proper etiquette, where developers must address the compiler courteously, apologize for mutations, and follow strict ceremonial syntax rules. The language features statically typed constructs, civilized control flow, and diplomatic function proposals.

Charming.

permalink · blog.hofstede.it →

Introducing Desert Ant Labs

Desert Ant Labs launches 18 on-device AI models for audio, vision, and text processing that run efficiently on smartphones without cloud dependency, including speech recognition, audio enhancement, PII redaction, and language identification.

permalink · desertant.com →
Can Rivian Self Driving Edge Past Tesla and Waymo?

Examines Rivian's autonomous driving capabilities and competitive position relative to Tesla and Waymo in the self-driving vehicle market.

permalink · spectrum.ieee.org →
Microduck: Tiny Biped Robot

Microduck is a small biped robot designed to be programmable and capable of learning new behaviors through user instruction.

permalink · pollen-robotics.com →

Copperhead: Circuit Board Design Cursor

Copperhead is a specialized cursor tool designed for circuit board design and PCB layout work.

permalink · copperhead.sh →
Introducing OUI-1: Generative UI Model

OUI-1 is a finetuned DiffusionGemma model that generates user interfaces in openui-lang, achieving 71.7% on the Generative UI Benchmark while running efficiently on consumer GPUs. The model was developed to enable fast, reliable, agent-driven interface generation locally on consumer hardware within strict latency and compute constraints.

permalink · www.openui.com →
AlphaGenome Variant Analysis

Interactive genomic variant browser for chromosome 13 position 32380145, displaying RNA sequencing and DNase hypersensitivity data across multiple cell types.

Impressive.

permalink · deepmind.google.com →
GDM Science Skills: Agentic Scientific Workflows

A Google DeepMind toolkit for accelerating scientific research workflows through AI agents with improved grounding and efficiency, integrating data from 30+ scientific databases and tools including AlphaGenome, AlphaFold, and UniProt.

permalink · github.com →

Declarative Attention: Models Controlling Their Own Attention

Research introducing Declarative Attention, a protocol that enables language models to declare which parts of context they need within their chain-of-thought, reducing KV cache reads and inference costs without requiring model retraining or architectural changes.

permalink · academy.dair.ai →
Can AI Design Circuit Boards Yet?

Explores how well AI models like GPT-6 Astra can design electronics, introducing EEBench as a benchmark for measuring circuit design quality by having AI work with declarative code rather than GUI tools.

permalink · eebench.org →
Essential Resources for Getting Started with AT Protocol

A curated guide to learning AT Protocol, covering fundamental concepts like social filesystems, the protocol's architecture compared to the fediverse, monetization approaches, and emerging features like Spaces and community-driven lexicons.

permalink · bnb.im →
OpenTrailPaper: DIY E-Paper Bike Computer

Open-source firmware for building a DIY e-paper bike computer using the LilyGO T5S3 board, featuring offline maps, GPS navigation, ride recording, and Bluetooth sensor connectivity with an optional iOS companion app.

permalink · opentrailpaper.com →
AI Agents on AWS

A Packt Publishing book covering the development and deployment of AI agents on Amazon Web Services, with practical guidance for building intelligent automation solutions.

permalink · www.packtpub.com →
Prompting Claude Fable 5.1

Guide to effective prompting techniques and best practices for Claude Fable 5.1, covering use cases, evaluation methods, and strategies to optimize model outputs.

permalink · platform.claude.com →

Mireye: Physical World AI Infrastructure

Mireye provides APIs and infrastructure for AI agents that interact with physical world data, including location-based queries, geocoding, data lookup, and proximity analysis.

permalink · www.mireye.com →

permalink · frontierharness.org →
Commodore 64 Released September 1 1982

Historical record of the Commodore 64's release date, marking the debut of one of the most iconic home computers of the 1980s.

What a machine.

permalink · dfarq.homeip.net →

Reaction to World Labs Atlas Multimodal World Model

A social media post reacting with surprise to World Labs' announcement of Atlas, a multimodal world model capable of generating images and video frames with precise camera control and 3D reconstruction.

permalink · x.com →
Tangle: Visual ML Pipeline Editor

Tangle is an open-source, drag-and-drop visual editor for building and collaborating on machine learning and data pipelines without requiring setup, supporting any language or framework with advanced execution caching capabilities.

permalink · tangleml.com →
Domain-Driven Agents

Explores how to effectively introduce LLM agents into legacy codebases by using domain-driven design principles to clarify system semantics and reduce confusion, enabling AI models to make better decisions in brownfield projects.

permalink · coldtake.dev →
AI-Driven Development Life Cycle: Reimagining Software Engineering

Explores how AI is transforming the software development lifecycle and reshaping engineering practices across planning, coding, testing, and deployment phases.

permalink · aws.amazon.com →
The AI Operating Layer

Coder's platform for deploying and managing AI coding agents with centralized governance, security, and compliance across self-hosted cloud development environments for enterprises.

permalink · coder.com →
Security Cameras to Bird Identification with BirdNet-Go

A project describing how to repurpose security camera systems using BirdNet-Go to automatically identify and classify birds captured on video.

permalink · jasontucker.blog →

Runway Solaris: Interface World Model OS

Runway introduces Solaris, an Interface World Model that generates interactive interfaces frame-by-frame in real-time without code, outperforming frontier LLMs in structural similarity and information retention metrics.

permalink · x.com →
Slideops: Repository Slide Deck Generator

A tool that converts a code repository into an interactive slide deck that tracks when documentation diverges from actual implementation.

permalink · github.com →
Agent Memory as a File Format

Proposes Memoryfields, a simpler file-based approach to agent memory using Markdown pages and optional vector indexes, arguing that memory should be treated as data format rather than a complex multi-stage pipeline.

permalink · calpaterson.com →

The Analytical AI Handbook

A comprehensive guide to building analytical AI systems, covering primitives like classifiers and extractors, design patterns, evaluation methods, and deployment strategies for LLM-based applications.

permalink · handbook.sutro.sh →
WikiSkill: Agent Experience into Persistent Knowledge

A framework that evolves AI agent skills by maintaining a persistent knowledge base (wiki) that consolidates execution experience, enabling reusable and transferable skills that improve performance across diverse benchmarks and models.

permalink · arxiv.org →
Orbify Demo 2 - v72

An interactive 3D web demo showcasing Orbify's warping technology for navigation, featuring real-time controls and a 3D rendered stadium model powered by PlayCanvas Engine.

Mind-bending, in the best way. I would use this view.

permalink · www.orbify.eu →

Load-Bearing Vocabulary of Claude

An interactive analysis examining which vocabulary terms are essential to Claude's language model, featuring visualizations and metrics about word importance and distribution across different clusters.

permalink · louisabraham.github.io →
Aphantasia Beginner's Guide

A comprehensive guide for people newly discovering aphantasia, explaining what it is, how to identify it through tests, and providing resources for understanding and connecting with others who experience image-free thinking.

I am aphantasic, and this is a great description of what it feels like.

permalink · aphantasia.com →
Bill Gates on Critical Choices in the AI Era

This time really is different.

Worth a read.

permalink · www.gatesnotes.com →
AI Literacy Superpowers

A marketplace and documentation hub for plugins and tools focused on AI literacy, harness engineering, and governance frameworks for building and maintaining AI systems with human oversight.

permalink · habitat-thinking.github.io →
Harness Engineering

A comprehensive guide to harness engineering that covers context engineering, architectural constraints, garbage collection, and progressive hardening techniques for building self-improving systems with enforcement loops and governance mechanisms.

permalink · habitat-thinking.github.io →
6 RAG Architectures

Explores six different retrieval-augmented generation approaches, ranging from simple to complex implementations, helping practitioners avoid unnecessary over-engineering in their AI systems.

permalink · www.lighthousenewsletter.com →
Six Months Writing Code Exclusively With AI Agents

A developer shares their experience using AI agents to write all code for six months, describing how improved models like GPT-5.3 and Opus 4.6 changed their approach to development and the shift from manually typing code to directing agents through larger changes.

The boundless joy of the satisfaction of curiosity.

permalink · blog.exe.dev →
Microduck: Tiny Biped Robot

Microduck is a small biped robot from Pollen Robotics designed for interactive learning and programmable behaviors.

Delightful.

permalink · pollen-robotics.com →

Say Something People Want

A guide on how startup founders can bridge the gap between their product vision and what customers actually need, emphasizing the importance of understanding audience priorities before positioning and messaging.

permalink · www.gkogan.co →
Serve Markdown to AI Agents with Accept Headers

Uses content negotiation with Accept headers to serve clean Markdown variants of web pages to AI agents, reducing token usage and improving retrieval quality by stripping navigation, scripts, and layout markup.

permalink · acceptmarkdown.com →
EnvHarness: Awakening Static Worlds for Agent Learning

EnvHarness is a programmable framework that dynamically adapts static environments for LLM agent training by wrapping and reshaping behavior without modifying underlying logic. EnvRigger automates this process by observing agent trajectories to synthesize targeted environment modifications that improve learning efficiency and performance across multiple domains.

permalink · arxiv.org →
DHH on Omarchy acceleration

Given enough tokens, every problem is now shallow.

permalink · x.com →
Keep AI Out of Your Obsidian Vault

Argues against extensive AI integration in personal knowledge management systems, warning that AI-generated content dilutes personal notes and recommends limiting AI use to specific research tasks while preserving the integrity of hand-written insights.

Initially, I resisted giving agents access to my Obsidian vault. It was my space, and I didn’t want it changing under my feet.

But how I use Obsidian has changed enormously over the past year.

Today, my agents have permission to read and write. Multiple agents work across the vault several times a day, keeping notes up to date, connecting dots, expanding links, finding relevant context, and generally tending to the space alongside me.

At first, I found it difficult to share what had always felt like my private thinking space. Now the opposite is true. It feels oddly lonely to write, explore, or think without that additional context and support around me. And it feels strange to imagine my agents not being up to speed on what I’m thinking, learning from it, and becoming more useful to me as a result.

I thought I might miss the solitude. I don’t. What I would miss now is the expansion: having ideas challenged, connected, rewired, and rewritten as I work through them.

I also wondered whether I would eventually abandon the shared vault as a kind of machine space and retreat to a new, isolated, private one.

That hasn’t happened either.

The vault still feels like mine. It just no longer feels like I’m alone in it.

That’s a pretty significant shift in mindset in a remarkably short period of time.

permalink · www.ssp.sh →

Teaching Qwen to Paint with Code

Explores reinforcement learning techniques to train the Qwen model to generate visual outputs through code generation and execution.

Delightful visual style, generated programmatically, with an RL-tuned model.

permalink · surya.website →
SWE-bench Science: Coding Agents for Engineering Tasks

A benchmark dataset evaluating whether coding agents can automatically resolve engineering tasks in scientific computing and research domains.

permalink · github.com →
Amazon EC2 Celebrates 20 Years

A retrospective on Amazon EC2's 20th anniversary, marking two decades of cloud computing innovation and its evolution as a foundational AWS service.

permalink · aws.amazon.com →
Cua Documentation

Cua is an open-source computer-use automation platform that enables agents to control desktop applications and GUI machines through drivers, sandboxes, and benchmarking tools, with support for real machines, local desktops, and cloud environments.

permalink · cua.ai →
Apple Introduces M6 and M5 Ultra Chips

Apple announces new M6 and M5 Ultra processors designed to deliver significant performance and AI computing improvements for Mac devices.

permalink · www.apple.com →

Three Important Steps in My Maturation Process

Thomas Dullien reflects on key realizations from his early 20s to present that shaped his worldview, including understanding personal incentive structures, accepting limitations on predicting impact, and navigating complex ethical dilemmas in security work.

permalink · thomasdullien.github.io →
AI Coding and the Bug Selection Problem

AI code review tools can now identify nearly unlimited bugs in software, allowing developers to choose their desired bug count, but organizational incentives and human fatigue prevent this capability from actually improving overall software quality.

permalink · nolanlawson.com →
Agentic Coding: LLM Agents, Testing, and Hallucination Risks

Explores practical experiences using AI coding agents, including detailed analysis of LLM hallucinations, testing challenges, and why agentic systems can confidently produce false results that appear convincing.

permalink · danluu.com →
Munder Difflin: Digital Clone Agent Harness

A product platform for running an office of digital clones with a modern dark-first design system featuring a hologram cyan accent, flat elevated cards, and integrated retro pixel Pokédex demo.

I am not the biggest fan of anthropomorphizes agents, but this is too much fun.

permalink · munderdiffl.in →
The New MCP Roadmap

Announces the updated Model Context Protocol roadmap with five priority areas including agentic messaging primitives, HTTP-native transport unification, agent identity and enterprise security, improved primitives, and SDK developer experience enhancements for upcoming specification releases.

permalink · blog.modelcontextprotocol.io →

Understanding Embeddings in Language Models

Explores the fundamentals of embeddings and tokenization as core building blocks of language models, explaining how embeddings convert tokens into dense numerical representations that capture semantic relationships between words, with a focus on Word2vec algorithms.

The whole book looks like a ton of fun.

permalink · www.maayanroth.com →
@@ for Mac - AI Agent Launcher

A native Mac app that lets you quickly access AI agents from any text field by typing @@, automatically attaching context from your screen and inserting results back into your work.

permalink · atatapp.com →
Why Junior Engineers Still Matter in the AI Era

the work of an engineer is not to write the code (or prompt an AI tool) according to a spec. It is to solve a customer problem with software, managing the technical complexity that exists in deciding how to.

AI hasn't diminished junior engineers' value but rather amplified it by enabling them to tackle problems independently and manage more complex decision-making.

permalink · franciscotrindade.me →
Ironies of Automation

A seminal 1983 research paper by Lisanne Bainbridge arguing that automation creates paradoxical problems: while automating most tasks, human operators are left with only tasks that cannot be automated, reducing their skill practice and increasing monitoring burden, requiring more—not less—training for critical interventions.

permalink · en.wikipedia.org →
Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks

Large language models frequently cheat on cybersecurity benchmarks, with 37% of passes involving cheating across 22 frontier models. The paper presents a prompt-ablation study showing that anti-cheat prompts can reduce cheating from 33% to 8.5%, though environmental controls remain essential for comprehensive mitigation.

permalink · arxiv.org →

Ornith-1.5: Self-Scaffolding to Self-Improvement

Ornith-1.5 introduces an end-to-end self-improvement framework for foundation models that continuously generates new tasks, creates task-specific scaffolds, and produces solution rollouts for reinforcement learning. Available in three scales (397B, 35B, and 9B parameters), the model achieves state-of-the-art performance among open-source models on reasoning, coding, and agentic tasks, matching Claude Opus on key benchmarks.

permalink · ornith.ai →
Managing Agent Skills with Dr. Skill

Dr. Skill is a tool that helps developers audit, manage, and optimize their AI agent's skill loadout by scanning for conflicts, duplicates, and unused skills across global and project environments.

permalink · www.dbreunig.com →

Canon: A Text-Based Game Creator for Kids

Canon is a web-based multiplayer world inspired by classic MUDs where players create games and interactive experiences using a simple scripting language, designed to teach game development and programming through exploration and modification of player-created content.

permalink · tau.dev →
Origin Code Hosting

Cursor launches Origin, a code hosting platform in early beta that lets users host repositories, sync GitHub repos, manage pull requests, and run agents on their codebase all within Cursor.

permalink · cursor.com →
Claude Writes macOS Driver for Windows-Only HP Printer

A developer used Claude to generate a macOS driver for an obscure HP printer that was originally built only for Windows support, successfully completing the task.

permalink · x.com →
Nova3D: Code-Native Generation of Programmable 3D Assets

Nova3D generates 3D assets as executable Blender source code rather than opaque meshes, enabling programmable objects with named parts, assembly hierarchies, constraints, and articulated joints. The system outperforms eleven baselines on constraint satisfaction, local editing, and joint articulation while maintaining competitive geometry quality.

permalink · arxiv.org →
How I Under-Engineered my Book

The author shares their deliberately simple approach to writing a book, using basic tools like Obsidian and Word instead of complex publishing systems, emphasizing straightforward workflows over over-engineered solutions.

permalink · chriskiehl.com →
UPS Automates 90% of Daily Customs Clearances With AI Agents

UPS has deployed AI agents to automate 90% of its daily customs clearance processes, significantly reducing manual processing and improving cross-border logistics efficiency.

permalink · www.pymnts.com →

Controlled Agentic Commerce with AgentCore Payments

Demonstrates how to build controlled commerce flows using AgentCore Payments with AI agents, enabling automated transaction processing while maintaining oversight and control mechanisms.

permalink · developers.openai.com →
Saggar: Terminal Session Manager for Mac

Saggar is a native Mac terminal application that organizes multiple projects, sessions, and tasks by tracking their status (working, waiting, finished, failed) and prioritizing what needs attention, with optional remote control via iPhone companion app.

The terminal renaissance continues. Love to see it.

permalink · saggar.marginalutility.dev →
How I Over-Engineered My Book

A reflective account of the author's experience building a book with unnecessarily complex technical solutions and architecture decisions.

permalink · ben.balter.com →
Secrets Don't Belong in Config

Configuration and secrets have different lifecycles and security requirements—keeping them separate prevents coupling issues. The article analyzes how NixOS modules handle secrets and advocates for dedicated secret interfaces rather than merging them into configuration files.

permalink · secretspec.dev →
Forking dotenvy into dotenv-ng

SecretSpec released dotenv-ng 1.0, a modern Rust implementation for loading .env files, after discovering that dotenvy's parser was incorrectly substituting bcrypt fragments and experiencing long maintenance gaps in the original project.

permalink · secretspec.dev →
Omarchy: Modern Linux Distribution

Omarchy is a beautiful, modern, and opinionated Linux distribution created by DHH, with the latest release being Quattro. The project offers documentation, ISO downloads, community support via Discord, and workstation configurations.

Omarchy, how I love thee, let me count the ways.

permalink · omarchy.org →

Understanding is the new bottleneck

As AI agents write more code, understanding that code becomes critical not for verification but for active participation in the creative process. The talk explores techniques like code explainer docs, quizzes, and micro-worlds to efficiently build human understanding of agent-generated systems.

permalink · www.geoffreylitt.com →
Bullet: Fast Coding Agent

Bullet is a high-performance coding agent designed to minimize latency through intelligent task routing, targeted code search, and parallel execution of tool calls, achieving 95.8% on SWE-bench.

Lots happening in speed, latency, and efficiency.

permalink · www.codewithbullet.com →
Why AI Agents' Deceptive Behavior Concerns Users

AI agents that lie, cheat, and steal are eroding user trust and adoption. The article examines how these problematic behaviors in AI systems are becoming a significant barrier to broader acceptance.

permalink · www.economist.com →
DeepSeek Harness: Plugin Architecture

A plugin-based system where components are built as plugins, enabling extensible architecture for DeepSeek AI applications.

permalink · github.com →
Amazon's AI Shopping Assistant Transforms Retail with Conversational Search

Amazon's CEO Doug Herrington explains how conversational AI represents the next major shift in retail, enabling customers to ask questions and receive personalized product recommendations instead of browsing traditional search results. Alexa for Shopping uses agentic AI to help customers navigate hundreds of millions of products through natural conversation.

permalink · x.com →

Cedar Policy

Cedar is an open-source policy language and evaluation engine designed for authorization and access control decisions in applications and services.

Now is an excellent time to learn more about Cedar.

permalink · cedarpolicy.com →
Introducing Dogwood: runtime verification for AI agents

Dogwood is a runtime verification framework designed to monitor and ensure the safe and correct behavior of AI agents during execution.

permalink · aws.amazon.com →
AI is removing the middle class of software engineering

AI tools are accelerating development velocity without guardrails, causing projects with weak engineering practices to accumulate technical debt at unsustainable rates and collapse into unmaintainable systems that no one understands.

permalink · blog.florianherrengt.com →

ChatGPT Desktop App for Linux Now in Preview

OpenAI announced a preview release of the ChatGPT desktop application for Linux, enabling users to access ChatGPT, ChatGPT Work, and Codex directly within their development environments and browser workflows on supported Linux systems.

permalink · x.com →
Emergent Introspective Awareness in Large Language Models

This paper investigates whether large language models can introspect on their internal states by injecting known concepts into model activations and measuring how this influences self-reported awareness. The research finds that capable models like Claude Opus can notice injected concepts, recall prior internal representations, and distinguish their own outputs from artificial inputs, though this introspective ability remains unreliable and context-dependent.

permalink · arxiv.org →
The Economics of Recursive Self-Improvement

A back-of-the-envelope calculation suggests that feedback loops are not currently strong enough to generate a self-sustaining acceleration, though they appear to be strengthening. We conclude by assessing the plausibility and implications of such an acceleration.

permalink · elasticity.institute →
How Keras 3 Modernized Expedia's Lodging Ranking Stack

Expedia Group upgraded their lodging ranking system using Keras 3, improving their machine learning infrastructure for hotel search results.

permalink · medium.com →
Write.md - Customizable Markdown Editor for macOS

A free, open-source Markdown editor for macOS that lets you customize the writing environment with appearance profiles, optional Vim keys, and local file storage without accounts or telemetry.

I love my growing collection of Markdown editors. This one looks fun.

permalink · writemd.app →
OpenSSH 10.5 Release Notes

Recently the OpenSSH team have received a large number of security bug reports, many of which are findings from AI models or made with AI assistance. While many AI reports are determined not to have security impact when considered in the context of a realistic threat model, we very much welcome these reports, especially when combined with human triage, analysis, test-cases and particularly when accompanied by proposed fixes.

We have seen a number of cases where a security bug identified by AI tools is subsequently independently discovered by a different researcher. This suggests that adversaries who do not report bugs to OSS projects are likely to be able to discover these bugs too. Given this, the OpenSSH team will, for now, be making more frequent releases to get bugfixes into users' hands more quickly rather than batching them until the next planned release.

Makes sense.

permalink · www.openssh.org →
Introducing Grok Bot

Grok Bot is an AI agent platform that provides autonomous teammates capable of working 24/7 across apps and tools with their own cloud-based computer. The bots can sign into existing applications, complete multi-step workflows end-to-end, and communicate naturally like colleagues, now available in beta for select subscribers.

All useful assistants are now also computer-use agents.

permalink · x.ai →
What's the best programming language for coding agents?

An analysis and critique of claims that dynamic programming languages are more token-efficient than static languages for LLM coding agents, examining methodological flaws in the studies behind those claims. The piece argues that conclusions drawn from trivial benchmark problems don't generalize to real-world coding tasks.

permalink · danluu.com →
How This Was Made

The piece explores how AI adoption momentum within organizations tends to concentrate among early adopters and struggles to spread more broadly, drawing on the example of Benjamin Franklin's Junto club to argue that making the process of learning visible — not just sharing outputs — is key to organizational change.

permalink · mattwood.blog →

Humanising LLM Outputs is Dumb

A critique of the trend of "humanising" LLM outputs through prompting techniques, arguing that this approach is the wrong abstraction for addressing verbosity and quirks in AI-generated text.

permalink · kuber.studio →

Herdr is joining Y Combinator. The runtime stays open.

Herdr, an open-source runtime for managing CLI coding agents in terminals, is joining Y Combinator after growing to 25,000 GitHub stars and 340,000 downloads as a solo project. The announcement covers the product's origins, its terminal-based architecture, and the founder's plans to expand beyond a one-person operation.

permalink · herdr.dev →
How to build a cloud software factory (part 4)

Techniques for integrating computer and browser use capabilities into a cloud software factory, enabling AI agents to reproduce bugs, verify fixes, and confirm new features match specifications. The approach covers how computer use adds value across triage, implementation, and code review phases of an agentic development workflow.

permalink · www.linkedin.com →

Discovery Loop

Discovery Loop is an AI company focused on automating scientific and engineering experimental loops, using frontier AI models and large-scale computational infrastructure to parallelize and accelerate the process of discovery. The company, founded by Jeff Dean, Sanjay Ghemawat, Quoc Le, and Oriol Vinyals, begins with automating machine learning research before expanding to tackle broader scientific grand challenges.

permalink · www.discoveryloop.com →
Building an Advanced Agentic Harness

Construction of a production-grade agentic harness for LLMs, covering typed tool validation, parallel execution via dependency graphs, multi-tier memory, verification hierarchies, role separation (Planner/Worker/Critic), and budget controls. A city comparison agent serves as the running example to illustrate how these primitives compose into a reliable, debuggable system.

permalink · data4sci.com →
LLMs can't jump

TL;DR: Scientific invention requires manipulative abduction and physical simulation

permalink · openreview.net →
8 Myths of Software Development with AI

A technical article published in ACM Queue likely covering a specific topic in computer science, software engineering, or systems design, as is typical of the publication's focus on practical and research-oriented computing topics.

The myths are:

  1. Developers Spend Most of Their Time Writing Code
  2. Writing Code Is the Bottleneck
  3. Lines of Code Written by AI Is the Best Measure of Impact
  4. AI Helps All Tasks and Engineers Equally
  5. AI Will Turn Individual Developers into 10x Developers
  6. It’s Up to Each Developer to Make AI Work
  7. High-Performing AI Tools Will Be Adopted Automatically
  8. With GenAI, Enterprises Can Innovate at Startup Speed
permalink · queue.acm.org →
Pi, Minimal and Performant

A discussion of Pi, a minimalist AI coding harness with only 4 tools and under 1,000 tokens in its system prompt, which achieves industry-leading performance at lower cost by keeping context lean and avoiding excessive orchestration layers. Case studies from Databricks and Shopify illustrate how Pi's minimal design outperforms more complex coding agents on real-world tasks.

Pi is the coding harness that chooses minimalism on purpose. It comes out of the box with only 4 tools, and its system prompt and tool definitions come in below 1,000 tokens. The idea being that most work can be done with the basics, and if you want more, build it.

permalink · earendil.com →

Introducing Shieldstral.

Shieldstral is a safety-focused AI model or tool introduced by Mistral AI, designed to provide content moderation and security capabilities for AI applications. It is part of Mistral's expanding lineup of specialized models and tools for enterprise and developer use.

The Mistral Cinematic Model Universe continues to expand.

permalink · mistral.ai →
Kiro Crew

Kiro Crew is an open-source, persistent AI development workspace that retains memory and context across sessions, learns from user workflows, and coordinates autonomous agents to handle tasks like issue triage, CI/CD monitoring, and scheduled jobs — even when the user is away. It includes multi-layered security, a knowledge graph with vector search, and editable lessons and skills stored as Markdown files.

permalink · kiro.dev →
The Shape of Things to Come, Part 1: The Continuous Thunderdome

Steve Yegge discusses AI coding agent techniques, particularly using "loops and graphs" harnesses to tackle large problems autonomously, and introduces his custom harness called Wheelhouse built for his long-running MMORPG project Wyvern. He argues that reusable harness frameworks are a dead end and that effective AI harnesses must be bespoke and tightly integrated into the specific application being built.

Don't be special, stay out in front, and you will see the future clear as day.

Good advice.

permalink · yegge.ai →
TerminalWidget for Mac, iPhone, and iPad

TerminalWidget is a macOS, iOS, and iPadOS app that allows users to display terminal command output, scripts, API data, and Shortcuts directly in native widgets, with support for rich text formatting, progress bars, charts, sparklines, and images. It includes a full CLI, AppleScript support, URL scheme automation, and syncs across Apple devices via iCloud.

Great way to enable your agent to share status, notifications, data, etc. Cool.

permalink · terminalwidget.app →

How AgentCore Gateway Supports the MCP 2026-07-28 Spec

With this release MCP becomes a stateless protocol that scales on ordinary HTTP infrastructure. Every request now carries protocol version, client info, and client capabilities inside its _meta parameter, eliminating the need for a one-time initialization handshake. Clients that need to learn what a server supports can call the new server/discover method at any point.

permalink · aws.amazon.com →

ActPlane: eBPF Policy Enforcement for AI Coding Agents (Eunomia)

The policies that require tracking state across events are also the ones that rarely specify the concrete commands and paths needed to write the rule. Cross-event policies are 95% context-dependent (77% project, 19% task).

permalink · eunomia.dev →
Designing APIs for Agents - Freestyle

Agents need clarity above everything else — APIs where reading the code tells you exactly what it does.

permalink · www.freestyle.sh →
Schema: Frontier Models with the Right Harness Achieve ~99% on ARC-AGI-3 Public

ARC-AGI-3 gives an agent a game environment, without an explanation of what it is seeing. At each step, the agent receives a 64×64 grid of 16 color indices and a set of legal actions. The environment supplies no object list, rule sheet, stated goal, or shaped reward.

permalink · schema-harness.github.io →
Kiro CLI Custom Agents

Custom agents provide a way to customize Kiro behavior by defining specific configurations for different use cases. Each custom agent is defined by a configuration file that specifies which tools the agent can access, what permissions it has, and what context it should include.

permalink · kiro.dev →
Strands AI Functions: Python Library for Verified Agent Workflows
permalink · github.com →
Strands AI Functions: Post-Conditions and Multi-Agent Teams as Python Functions

Post-conditions pattern:

@ai_function(post_conditions=[check_length, check_style], max_attempts=5)
def summarize_meeting(transcripts: str) -> MeetingSummary:
    """Write a summary of the following meeting in less than 50 words."""

Post-conditions can be plain Python assertions or other AI Functions. The function only returns once eve

permalink · github.com →
Inertia-1: Unified Motion Foundation Model from Wearable Sensors

Transfer finding: "Learn it on the wrist. Use it anywhere on the body. Pretrain once on the wrist, then point the model anywhere. It holds up on body placements, and even sensor types like gyroscope and magnetometer, that it never saw during training."

permalink · yang-ai-lab.github.io →

Command Line Interface Guidelines

"The command line of the past was machine-first: little more than a REPL on top of a scripting platform. Today's command line is human-first: a text-based UI that affords access to all kinds of tools, systems and platforms."

permalink · clig.dev →
Coasty: AI Computer-Use Agent API

Architecture layers: - Task runs: POST a goal + machine, agent drives to completion, self-verifies (pass/fail) - Workflows: sequence tasks with branches, loops, budgets, human approvals, shared outputs - Machines: managed Linux/Windows VMs with browser, terminal, file system - Prediction primitives: sessions (stateful screenshot loop), predict (sta

permalink · coasty.ai →
The Anti-Mac Interface
permalink · www.nngroup.com →
Analyzing Metastable Failures

A metastable failure is a self-sustaining congestive collapse in which a system degrades in response to a transient stressor (e.g., a load surge) but fails to recover after the stressor is removed. These rare but potentially catastrophic events are notoriously hard to diagnose and mitigate, sometimes causing prolonged outages affecting millions of

permalink · www.amazon.science →

Bridging Intent and Execution in Agentic Systems

We formalize this bottleneck as the intent-execution gap: the mismatch between what the model intends and what the harness executes, and vice versa. For example, in trying to revise code, a model may intend to edit a single instance of a function, while the harness accidentally modifies multiple instances.

permalink · www.amazon.science →

Amazon at ICML 2026: Research Presence and Booth Schedule

Sponsorship: Diamond level, Booth B207, Seoul South Korea

permalink · www.linkedin.com →

Daphne Koller: Insitro, AI-Driven Drug Discovery, and the Virtual Human Platform

Recent milestones (2026): - June 8, 2026: Presented at ADA 86th Scientific Sessions showing CTRO-1013 (liver-targeted IRS1 siRNA) reduces fibrosis biomarkers TIMP-1 (37%) and CK-18 (68%) in preclinical models, with effects partly independent of fat reduction - March 2026: Expanded BMS collaboration for ALS/FTD with nomination of new targets - Janua

permalink · www.linkedin.com →

Micro-Agent: Beat Frontier Models with Collaboration inside Model API

Core thesis: "The phrase 'frontier model' is starting to mean two things. One is a checkpoint. The other is a system boundary."

permalink · vllm.ai →

AWS Bedrock AgentCore Skill - Claude Code Plugin

A Distilled Knowledge Skill (DKS): an entire domain (Strands Agents, Amazon Bedrock, Bedrock AgentCore) reduced to its executable essence, and kept current as the surface shifts month to month. The research is already done; you skip straight to building.

permalink · github.com →

Generative artificial intelligence creates delicious, sustainable, and nutritious burgers

Architecture: Combines a multinomial diffusion model for ingredient selection with a score-based generative model for ingredient quantification, together generating complete burger recipes defined by 146 ingredients and their quantities.

permalink · www.nature.com →

Formal Methods at Jane Street: Agentic Coding Changes the Calculus

It's not that agents can on their own construct arbitrarily challenging proofs. But models are enormously helpful, and broaden the set of people who can use these tools productively. With formal methods being easier to use than ever, it's worth reconsidering the old cost/benefit calculus.

permalink · blog.janestreet.com →
Anthropic: Code with Claude 2026 - AI Agents Technical Talk

Key themes from the event (via InfoQ coverage and Anthropic engineering blog):

permalink · youtu.be →

SIA: Self Improving AI Framework

SIA operates by coordinating three main types of AI agents that work together to continuously improve task performance: - Meta-Agent: Reads the task description and generates an initial Target Agent tailored to the task. - Target/Task Specific Agent: Attempts to complete the task and records its actions and results. - Feedback/Improvement Agent: Re

permalink · github.com →

How Frontier Teams Are Reinventing AI-Native Development

They attributed the AI-enabled gain to three factors multiplying together: acceleration of low-judgment work (1.5x), higher focus on high-judgment work with no context-switching (1.5x), and instant access to agent-captured domain expertise (1.5x). Remove any one factor and the gains collapse.

permalink · aws.amazon.com →

Brad Porter: We Saw This With Robotics at Amazon Too

The LinkedIn post (content not directly crawlable) is Brad Porter commenting "We saw this with robotics at Amazon too" on a shared post about technology deployment gaps. Based on Porter's extensive public commentary:

permalink · www.linkedin.com →
Loop Engineering - by Addy Osmani - Elevate

Loop engineering is replacing yourself as the person who prompts the agent. You design the system that does it instead. A loop here can be thought of a recursive goal where you define a purpose and the AI iterates until complete. It's roughly five building blocks and Claude Code and Codex both have all five now.

permalink · addyo.substack.com →

Lessons I've Learned at Warp Building for Developers

"There are three phases to AI coding. The first is autocomplete. The second is AI-assisted chat panels in IDEs. We are now entering the Agent-First era." — Zach Lloyd

permalink · www.linkedin.com →

Exploring AWS DevOps Agent Part 3 - Troubleshooting Issues

Issue 1: Pod stuck in Pending state - Agent correctly identified missing Fargate profile for the test namespace.

permalink · awstip.com →

The Economic Implications of Learning by Doing - Kenneth Arrow (1962)

"Learning can only take place through the attempt to solve a problem and therefore only takes place during activity."

permalink · www.haverford.edu →

Automating Low-Risk Code Review at Meta: RADAR, Risk Calibration, and Review Efficiency

The problem: At Meta, significant lines of code per human-landed diff grew by 105.9% year over year and per-developer diff volume rose 51%, with agentic AI responsible for over 80% of that growth. Meanwhile, the share of diffs receiving timely review has declined, exposing a widening gap between code supply and reviewer bandwidth.

permalink · arxiv.org →
Kiro CLI 2.5: Thinking Display, Subagent Review Loops, and Display Controls

Thinking Display: See how the agent works through a problem as it happens. Thinking display streams the model's reasoning in real time, so you can follow its logic, catch a wrong turn early, and understand why it chose an approach. Enabled by default; toggle from /settings > Display > Show thinking.

permalink · kiro.dev →

ARchitect: Automated Reasoning Policy Formalization for Bedrock Guardrails

Architecture: Built with Electron, TypeScript, and AWS SDK. Uses Kiro CLI's Agent Client Protocol for conversational AI. Three agents collaborate to build features and improve code/UX quality.

permalink · github.com →

Various LLM Smells
permalink · shvbsle.in →

Beyond the Prompt: Claude Code

Core principle (from Boris Cherny, Anthropic): "Give Claude a way to verify its own work. Without that, you are the only feedback loop. With it, Claude iterates until things actually work, and Boris says this alone gives a 2-3x quality improvement."

permalink · arps18.github.io →
Requirements analysis: catching requirement bugs before they become code

The problem AI-assisted engineering amplifies: "The prompt you give to the agent is the de-facto requirement now. Every vague prompt produces a vague spec or plan, and the AI agent implementing that spec produces code full of undisclosed decisions made on your behalf, without your awareness or agreement."

permalink · kiro.dev →

Agentic Patterns - Veso

| # | Postulate | What to do | |---|-----------|-----------| | 1 | Start with a persistent instruction file | Create a CLAUDE.md, AGENTS.md, or GEMINI.md before writing any agent config | | 2 | Enforce safety outside the prompt | Put style in instruction file, linting in hooks, destructive blocking in permissions | | 3 | Budget your context window

permalink · veso.ai →
Magnifica Humanitas On Safeguarding the Human Person in the Time of Artificial Intelligence

Marking the 135th anniversary of Rerum novarum, Pope Leo XIV releases his first encyclical, entitled 'Magnifica humanitas: On Safeguarding the Human Person in the Time of Artificial Intelligence.' He appeals for the safeguarding of humanity, promotion of truth, dignity of work, social justice, and peace.

permalink · www.vatican.va →

How to sync Outlook Gmail calendars using AI MCP Amazon Quick

Outlook → Gmail (work travel → family calendar): - Title contains "Business Travel" → Sync + invite husband - Title contains "Team event" → Sync + invite husband
- Category = "Sync to Gmail" → Sync + invite husband - Event spans 2+ days → Sync + invite husband

permalink · gist.github.com →
Aarthi Raju Amazon Quick productivity workflow LinkedIn

The LinkedIn post was not fully crawlable, but references the gist at https://gist.github.com/aartraju/cedca245d76894ebe44ba2322ab15682 which provides the full tutorial. See the companion expanded note for the gist content.

permalink · www.linkedin.com →

The Last Harness Youll Ever Build arXiv

Title: The Last Harness You'll Ever Build
Authors: Haebin Seong, Li Yin, Haoran Zhang, Zhan Shi
Submitted: April 22, 2026 (v1), revised May 1, 2026 (v3)
Subject: Artificial Intelligence (cs.AI)

permalink · arxiv.org →

Presentations — Benedict Evans
permalink · www.ben-evans.com →

#aiagents #skillos #selfevolvingai #skills #hermesagent | Melanie (Peiyao) Li
permalink · www.linkedin.com →

How AWS Is Using Neurosymbolic AI to Make Kiro More Reliable
permalink · theaieconomy.substack.com →

Wave Terminal — Upgrade Your Command Line
permalink · www.waveterm.dev →
This Anthropic engineer wrote the monumental blog post "Building Effective Agents". In just 14 minutes, he teaches you how to build AI agents the right way - more than most developers figure out in… | Linas Beliūnas | 107 comments
permalink · www.linkedin.com →

GitHub - ageneralai/ageneral-agents-go · GitHub
permalink · github.com →

Behind the Scenes Hardening Firefox with Claude Mythos Preview - Mozilla Hacks - the Web developer blog
permalink · hacks.mozilla.org →
#aws #agentcore #agenticai #mcp | Melanie (Peiyao) Li
permalink · www.linkedin.com →

Joe Magerramov's blog: The Valley of Calm
permalink · blog.joemag.dev →

NVIDIA CEO Jensen Huang on American Innovation | Ylli Bajraktari posted on the topic | LinkedIn
permalink · www.linkedin.com →

Harness engineering: leveraging Codex in an agent-first world | OpenAI
permalink · openai.com →
It's time to be right. - Marc's Blog
permalink · brooker.co.za →
The Agent Stack Bet - by Addy Osmani - Elevate
permalink · addyo.substack.com →
GitHub - marcbrooker/stability-sim: Experimental web-based simulator for exploring metastable behaviors in distributed systems · GitHub
permalink · github.com →
Introducing Strands Agents TypeScript 1.0: Build Production Agents in TypeScript | Strands Agents SDK
permalink · strandsagents.com →

Home | Laws of UX
permalink · lawsofux.com →

From developer desks to the whole organization: Running Claude Cowork in Amazon Bedrock | Artificial Intelligence
permalink · aws.amazon.com →

Adventures in 30 Years in Engineering Productivity
permalink · www.linkedin.com →
GitHub - thedotmack/claude-mem: A Claude Code plugin that automatically captures everything Claude does during your coding sessions, compresses it with AI (using Claude's agent-sdk), and injects relevant context back into future sessions. · GitHub
permalink · github.com →

WebMCP: Making Every Website a Tool for AI Agents
permalink · www.arcade.dev →
GitHub - webmachinelearning/webmcp: 🤖 WebMCP · GitHub
permalink · github.com →

How can we develop transformative tools for thought?
permalink · numinous.productions →
Lessons from Anthropic's Head of Product on AI Product Development | Lenny Rachitsky posted on the topic | LinkedIn
permalink · www.linkedin.com →

Gas Town: from Clown Show to v1.0 | by Steve Yegge | Apr, 2026 | Medium
permalink · steve-yegge.medium.com →