We use analytics and advertising tools by default. You can update this anytime.
Manage optional tracking categories. Necessary cookies stay on so the site can function.
Staff Writer & AI Editorial Lead
Katie Parrott is a staff writer. She writes Working Overtime and contributes to Vibe Checks, Source Code, and Context Window.
Off-the-shelf AI couldn’t meet Headway’s security and workflow needs, so it built its own. Now the whole company uses it.
Plus: What to do when your coding model disappears, GitHub's COO on 14 billion agent commits, and engineers questioning their future
Kieran Klaassen turned a prompt into a working app in an hour and shows you how
OpenAI's latest model update excels at instruction-following and extended tasks, but don't expect it to surprise you
Start with three simple tools, and let the AI figure out the rest
The app worked. Nobody, including me, had checked whether it was safe.
Plus: Vercel and Lovable’s security woes, and how to make your agent your watchdog
The asynchronous, agentic workflow developers love is finally accessible to everyone—but the polish isn't there yet
A mini-Vibe Check on Gas City, a Grok classifier that grades your X drafts, and why HTML is the new markdown
OpenAI nailed the interface. But it's built for hardcore engineering.
Anthropic's latest Opus is more precise, more literal, and the best coding model we've tested on well-specified tasks—but it won't fill in the gaps for you anymore
Plus: Agent-designed automations, why final review belongs in the destination app, and how to use our compound knowledge plugin
Plus: A manual for handing recurring work to Opus, a skill for judging new tools in context, and the specialist model that beat the frontier for less
Sonnet 4.6 delivers Opus-close performance at half the price—but speed didn't come along for the ride
Plus: Perplexity’s rules for agent skills, the office politics of dictation, and creating a weekend AI piano coach
Plus: the number KateBench taught us to count, and a six-agent solar crew helping decide whether to run the dryer
Plus: Dan’s attempt to clone Kate, a shortcut for turning demonstrations into skills, and the human goals machines still need us to set
Plus: The Vatican weighs in on AI labor, and our Codex playbook
Opus 4.8 tops both our Senior Engineer benchmark and our writing tests. It’s the most complete model we’ve tested. We just wish it had an app to match.
Plus: An AI writing policy that makes writers do the thinking, Zuckerberg’s AI Future for Everyone runs on Meta, and a tool that keeps your agents from timing out