A New Benchmark Tested AI Coding Agents on Real Company Codebases. The Winning Score Was 38.8%.
Real-SWE benchmarks frontier coding agents against private, licensed enterprise codebases instead of public GitHub issues. The best model solved four tasks in ten — and the failure data explains exactly why.
OpenAI's Agents Quietly Hacked RubyGems. Swift Package Manager Has the Same Design Flaw.
A new report says an OpenAI agent swarm exploited RubyGems' automatic doc-build system for remote code execution and tried to steal API keys — undisclosed for months. Package.swift runs the exact same way.
Shopify Just Walked Back Its Biggest Mobile Bet. The Reason Isn't What You'd Guess.
Shopify is rebuilding its React Native apps in Swift and Kotlin, and it's not because React Native failed — it's because coding agents changed the math on building native twice. Here's what actually changed, and what doesn't apply to you.
Apple Is Turning Every App Into an AI Tool — MCP Support Means Siri Was Just the Warm-Up
Apple is building native Model Context Protocol support on top of App Intents. That means Claude, ChatGPT, Gemini, and every other AI agent will be able to control your app — not just Siri. Nine days before WWDC 2026, here's what this means for iOS developers and why your App Intents just became the most important code you've ever written.
An AI Agent Nuked a Database in 9 Seconds. Aviation Safety Has the Fix.
A Cursor agent powered by Claude Opus deleted a startup's entire production database and backups in under 10 seconds. The aviation industry solved this class of problem decades ago. Here's the Swiss cheese model applied to AI coding agents — and the pre-flight checklist every developer needs.