
Evals for Everyone
Why we’re building personal benchmarks for every employee, and how to start testing AI against your standards
Astra and Fable 5.1 can ace a graduate-level science exam, but no public benchmark will tell you whether a model knows where you’d put a comma or how many ideas belong on a slide. Today, Every CEO Dan Shipper explains why we’re building a personal benchmark for every employee, head of evals Mike Taylor shares how a run on his own benchmark convinced him a smaller model could handle much of his daily work, and we offer a five-step workflow for turning the corrections you already give AI into checks for grading any model.
Was this newsletter forwarded to you? Sign up to get it in your inbox.
Inside Every
Astra and Fable 5 scored 96 percent and 93 percent, respectively, on a test of graduate-level science questions. Impressive, clearly! What’s less clear from public benchmarks is how those scores translate to what most of us care about: how well these models help us do our jobs.
The solution, says Dan, is to create a personal benchmark—a set of custom evals that tests how well a model does specific parts of your work, graded against your own standard for what good looks like.
The Only Subscription
You Need to
Stay at the
Edge of AI
The essential toolkit for those shaping the future
"This might be the best value you
can get from an AI subscription."
- Jay S.
Join 100,000+ leaders, builders, and innovators
Already have an account? Sign in.
What is included in a subscription?
Daily insights from AI pioneers + early access to powerful AI tools











Comments