
Benchmarks Don’t Know Your Job
Plus: the number KateBench taught us to count, and a six-agent solar crew helping decide whether to run the dryer
Was this newsletter forwarded to you? Sign up to get it in your inbox.
AI has a measurement problem. Companies know how much they spend on models and how those models score on public benchmarks. What they often don’t know is whether the models save employees time or produce work people can trust without rechecking. In today’s Context Window is a look at how companies can answer those questions for themselves: We explain why organizations need tests built around real work, show what cloning our editor in chief has taught Every about evals, and meet the six-agent crew helping an Every engineer decide whether his family can run the dryer.
Introducing Attio: the agentic CRM
Transform the way revenue work gets done with Attio. Get agents that build pipeline, convert leads, and run all your sales motions. Your agents track the whole book, so you save the ones slipping and grow the ones rising. Then Ask Attio any question about your business, from the weekly forecast to performance by rep, and get the answer in seconds.
For teams building the next era of revenue.
The Only Subscription
You Need to
Stay at the
Edge of AI
The essential toolkit for those shaping the future
"This might be the best value you
can get from an AI subscription."
- Jay S.
Join 100,000+ leaders, builders, and innovators
Already have an account? Sign in.
What is included in a subscription?
Daily insights from AI pioneers + early access to powerful AI tools












Comments