We use analytics and advertising tools by default. You can update this anytime.

Choose the best model for the work you do
Was this newsletter forwarded to you? Sign up to get it in your inbox.
Nobody hires a vice president based on their SAT scores. Yet every time a new model drops, AI researchers first check how well it does on a Math Olympiad and a set of multiple-choice trivia, then argue online about whether it’s the smartest model in the world. Wharton professor Ethan Mollick points out that Massive Multitask Language Understanding-Pro (MMLU-Pro), one of the most-cited benchmarks, asks models for the approximate cranial capacity of Homo erectus and the place named in the title of Cheap Trick’s 1979 live album. Those tests measure something—the scores are directionally useful—but not what you need to know: Can this model help you with your job? To answer that, you need to build a personal benchmark.
I joined Every earlier this year to run our technology consulting practice, so I spent most of my time getting models to write, build dashboards, and assemble slide decks. With each new model, I developed an instinct for when and how to use it. Claude Opus 5 felt argumentative. Claude Fable 5 felt like talking to a genius. GPT-5.6 Sol felt like a safe pair of hands. But when friends and colleagues asked me for recommendations, I couldn’t always defend my choices. It was also time-consuming to test every model on every type of task. I often missed opportunities to use something better or cheaper. So, inspired by Mollick’s piece about why you should test models on your own work instead of trusting benchmarks, I built a personal one: a small, private test that tells me whether I like working with a model, without spending all my time testing—because I have a job to do.
Now I’m head of evaluations (evals) at Every, testing new models from the frontier labs for the qualities we value in our work. I want to help everyone on the team build personal benchmarks for their own work. Building personal benchmarks has three levels, and each can make you more confident you’re using the right model for your job. If you set up your personal benchmarks correctly, you’ll know when to switch between Astra and Fable, and whether you can save money using a smaller, cheaper model like Luna or Haiku.
The essential toolkit for those shaping the future
"This might be the best value you
can get from an AI subscription."
- Jay S.
Join 100,000+ leaders, builders, and innovators

Already have an account? Sign in.
Daily insights from AI pioneers + early access to powerful AI tools
Comments