
The Model Is the Easy Part
Plus: Plus: Codex folds into ChatGPT, and why the next drug giant won't be an AI company
Was this newsletter forwarded to you? Sign up to get it in your inbox.
Measure what matters—and get paid for it
Token spend is climbing everywhere. Some of that AI use is valuable; some isn’t. The only way to tell the difference and spend efficiently is to measure—but measure what?
Getting AI tools into your team’s hands was part one of AI adoption. Today, companies must identify business-specific, measurable goals. That part, too, requires both a definition of “good” and the data to measure how close you’re getting to it. Meanwhile, frontier labs are under pressure to improve at economically valuable domains such as finance, life sciences, and general reasoning. They’re paying for data that helps them get there. Well-defined goals now create value for your company in two ways: return on investment for you and your customers and data that can help others improve, too.
Over the past year at Good Start Labs, we’ve built benchmarks, trained agents, and helped game publishers operationalize and monetize their data for that lab market. Arkadium is one publisher with hundreds of games played by tens of millions of players. We supported its recent launch of Game Lab, a public leaderboard scoring how well frontier models play simple games, in partnership with Meta and DeepMind. Arkadium set a clear goal: Give its players a good game against AI. Together we built the benchmarks and evaluated them against real users. Then the scores came in. The same models that make novel discoveries in math and science lose 90 percent of their Gin Rummy games—against casual players.
Create a free account, or log in.
Every members live and work at the edge of AI. Join now.
By continuing, you agree to the Terms of Sale, Terms of Service, and Privacy Policy.
Enjoy unlimited access to all of Every.
See subscription options