We use analytics and advertising tools by default. You can update this anytime.

Why we’re building personal benchmarks for every employee, and how to start testing AI against your standards
Astra and Fable 5.1 can ace a graduate-level science exam, but no public benchmark will tell you whether a model knows where you’d put a comma or how many ideas belong on a slide. Today, Every CEO Dan Shipper explains why we’re building a personal benchmark for every employee, head of evals Mike Taylor shares how a run on his own benchmark convinced him a smaller model could handle much of his daily work, and we offer a five-step workflow for turning the corrections you already give AI into checks for grading any model.
Was this newsletter forwarded to you? Sign up to get it in your inbox.
Astra and Fable 5 scored 96 percent and 93 percent, respectively, on a test of graduate-level science questions. Impressive, clearly! What’s less clear from public benchmarks is how those scores translate to what most of us care about: how well these models help us do our jobs.
The solution, says Dan, is to create a personal benchmark—a set of custom evals that tests how well a model does specific parts of your work, graded against your own standard for what good looks like.
The essential toolkit for those shaping the future
"This might be the best value you
can get from an AI subscription."
- Jay S.
Join 100,000+ leaders, builders, and innovators

Already have an account? Sign in.
Daily insights from AI pioneers + early access to powerful AI tools
Comments