
The Steering Wheel for AI
Plus: The slippery slope of vibe coding tools, and exploring machine creativity
Mar 23, 2025 · 9 min readUpdated Jan 17, 2026
Hello, and happy Sunday! As you rest and reflect on the week past and the week to come, we're thinking about AI benchmarks. We may think of benchmarks as a simple yardstick, but for today's models they are so much more—as our own Alex Duffy writes, they're a critical means of giving some direction to the wild AI ride we find ourselves on. Meanwhile, Katie Parrott wrote a fun first-hand chronicle of how vibe coding tools led her to want to learn to code on her own. And Rhea Purohit wrote a fascinating account of machine creativity that is guaranteed to stoke wonder, perhaps along with your own creative juices.—Michael Reilly
Was this newsletter forwarded to you? Sign up to get it in your inbox.
Benchmarks lead the way
For most of us, driving a car means harnessing a controlled explosion. You sit behind a masterfully engineered hunk of metal that turns burning gasoline into progress at 70 miles per hour. With hands lightly on a steering wheel and a foot on a pedal, you control incredible power.
Benchmarks steer AI the same way. AI is powerful—explosive, even—but without a clear sense of where you want to go, it’s easy to confuse activity with achievement.
Last week, Hugging Face shut down its famous Open LLM Leaderboard. For two years, the company evaluated more than 13,000 models, helping sort good from great. But as AI evolved, these benchmarks stopped measuring real-world impact. Models have been rapidly gaining new abilities—like reasoning and agents—that the leaderboard didn’t capture. Some teams were even training models for the express purpose of acing these benchmarks—essentially “training on the test,” which was no longer representative of real-world performance.
Make email your superpower
Not all emails are created equal—so why does our inbox treat them all the same? Cora is the most human way to email, turning your inbox into a story so you can focus on what matters and getting stuff done instead of on managing your inbox. Cora drafts responses to emails you need to respond to and briefs the rest.
This week, Metr re-captured the AI community’s attention with a new and well-chosen benchmark. The company's blog post showed off a clear demonstration of AI’s impact in striking terms.
Create a free account, or log in.
Every members live and work at the edge of AI. Join now.
By continuing, you agree to the Terms of Sale, Terms of Service, and Privacy Policy.
Enjoy unlimited access to all of Every.
See subscription options
