
The Model Is the Easy Part
Plus: Plus: Codex folds into ChatGPT, and why the next drug giant won't be an AI company
Was this newsletter forwarded to you? Sign up to get it in your inbox.
Measure what matters—and get paid for it
Token spend is climbing everywhere. Some of that AI use is valuable; some isn’t. The only way to tell the difference and spend efficiently is to measure—but measure what?
Getting AI tools into your team’s hands was part one of AI adoption. Today, companies must identify business-specific, measurable goals. That part, too, requires both a definition of “good” and the data to measure how close you’re getting to it. Meanwhile, frontier labs are under pressure to improve at economically valuable domains such as finance, life sciences, and general reasoning. They’re paying for data that helps them get there. Well-defined goals now create value for your company in two ways: return on investment for you and your customers and data that can help others improve, too.
Over the past year at Good Start Labs, we’ve built benchmarks, trained agents, and helped game publishers operationalize and monetize their data for that lab market. Arkadium is one publisher with hundreds of games played by tens of millions of players. We supported its recent launch of Game Lab, a public leaderboard scoring how well frontier models play simple games, in partnership with Meta and DeepMind. Arkadium set a clear goal: Give its players a good game against AI. Together we built the benchmarks and evaluated them against real users. Then the scores came in. The same models that make novel discoveries in math and science lose 90 percent of their Gin Rummy games—against casual players.
That’s because these models are shaped by what they’ve seen: lots of math, and almost no Gin Rummy. Popular AI models may not have experience with whatever you work on all day either. But defining the goal revealed options for Arkadium beyond a large language model—different forms of AI work well for different goals. Instead of an LLM, we trained an expert model with 4.6 million parameters in an 18-megabyte file that runs on a regular CPU. At full strength, it beats human players about 90 percent of the time, and the economics are as lopsided in its favor: At 1 million requests a day, our expert model would cost about $60 a year; a frontier LLM would run into the multi-millions.
For Arkadium, that “good game” goal paid twice: a better experience for its players and a new revenue line in the form of selling anonymous data to labs for millions. Frontier models performed poorly at games; Arkadium’s well-structured gameplay data from “good” games with “good” human players was what the labs needed to improve. As Arkadium continues to sell that anonymized data to frontier labs, its models will learn from it and apply that intelligence. (GPT-5.6 is already a much better cruciverbalist.)
Reddit, Shutterstock, and News Corp have already turned their data into hundreds of millions of dollars in recurring revenue. Most companies outside the labs’ competitive focus could benefit from selling them data. One major caveat: Companies whose data is their product—like Figma or Cursor—must protect that IP. Anthropic chief product officer Mike Krieger left Figma’s board days before Anthropic shipped a competing design tool; Figma CEO Dylan Field later said they were “not consistently candid.” Cursor built on Claude while Claude Code was described to it as a “research effort,” and then they watched it become a direct competitor. This week it was reported that Anthropic asked big pharma for their data, and nearly everyone said no. If the lab you’d sell to is entering your business, retaining a data edge is essential.
Nobody knows yet how defensible the data market is long term, only that the pot promises to be large. But choosing the right goal for AI, measuring progress, and valuing what you learn will only grow more important as AI becomes table stakes for most companies. Defining and measuring “good” is emerging as the next stage of AI adoption. It never ends—and is becoming a requirement to compete.—Alex Duffy
Was this newsletter forwarded to you? Sign up to get it in your inbox.
Knowledge base
“How I Polish Software That Agents Built” by Kieran Klaassen/Source Code: In compound engineering, agents now run most of the loop—planning, building, reviewing, and opening clean PRs overnight—so Kieran Klaassen’s work narrows to the last human step: polish, or judging whether what they shipped is good or simply functional. Read this for where human judgment sits in an agent-first workflow.
“The Case Against Skills” by Laura Entis/Context Window: Head of tech consulting Mike Taylor argues that most trending AI “skills” have become redundant now that frontier models absorbed them, and piling on instructions can make outputs worse and pricier. Also inside: Monologue general manager Naveen Naidu keeps just one skill, OpenClaw’s autoreview, which helped ship a four-app Monologue Notes feature in a single nine-hour overnight run, and head of growth Austin Tedesco has dropped skills entirely except compound engineering.
“The Urge to Merge (ChatGPT and Codex)” by Katie Parrott/Context Window: OpenAI folded its standalone Codex app into a new ChatGPT desktop with three modes—Chat, Work, and Codex—and power users revolted, with YouTuber Theo Browne calling it a “generational fumble.” Also inside: a workflow for putting a pricey model in charge of cheaper ones—ChatPRD founder Claire Vo treats Fable as a senior consultant, and Dan Shipper runs the same play across labs—plus a 53 percent drop in error rate when an agent saves reusable tools instead of rewriting code.
“The Ops Team That Routes Work Across Models” by Laura Entis/Context Window: Every’s business-operations team moves fluidly between Fable, Codex, and support platform Fin to hit deadlines that shouldn’t be possible. Executive operations manager Jalaiyah Bolden pointed Fable at scattered context and had a support plan, help articles, and 17 response templates ready for a product launch in about 90 minutes. Also inside: customer service manager Waqqas Mir turns a mishandled support chat into a single sharper agent instruction, and Spiral general manager Marcus Moretti spends two weeks with Sonnet 5.
🎧 🖥 “The Founder of a $1.5 Billion AI Company on What Comes After the First Wave of AI Apps” by Dan Shipper/AI & I: Granola cofounder and CEO Chris Pedregal joined Dan to explain why running a startup is “a knife fight” that never ends: After a $1.5 billion valuation on its AI meeting notetaker, Granola has watched Notion, OpenAI, and Zoom copy the feature. Watch or listen to this for why Granola is betting on owning the work around meetings, not just the notes. 🎧 🖥 Listen on Spotify or Apple Podcasts, watch on YouTube, or follow the discussion on X.
🖥 “I’m an Editor—And I Built Our Newest Feature” by Jack Cheng/On Every: Gift links let paid and All Access members share paywalled Every articles with anyone. Senior editor Jack Cheng, who isn’t an engineer, built the feature himself with Codex and current frontier models—a case study in how an AI-native company decides what to launch. Read this for how anyone on a 30-person team can now build and test a feature. 🖥 Watch Jack, Kate Lee, and Austin Tedesco discuss how gift links came to be.
From Every Studio
Every All Access is here
All Access, Every’s new annual membership, is built around a Builder Pack—more than $7,000 in credits and trials across 10 tools we use, at roughly 90 percent below market—plus unlimited accounts on Cora, Every’s AI email tool, unlimited usage on Spiral, Every’s writing tool, and everything in a paid Every membership.
A new product powered by Monologue
Sandbar’s Stream, a private voice ring that captures thoughts as spoken notes, is using the Monologue API to turn speech into text. It’s the first public product outside Every built on Monologue’s voice infrastructure.
Spiral expands Writing Rules
Spiral expanded its Writing Rules so more complex instructions about structure and phrasing can shape a draft. Alongside preferences such as avoiding em dashes or using Oxford commas, you can now ask Spiral to split, shorten, or reorganize copy in any language. The draft-count slider also reliably returns the number of options you choose, and a backend stability fix should reduce crashes.
Alignment
The narrow promise. One problem with so many companies hurtling into AI drug discovery is that finding better candidates faster and more cheaply is only one small part of bringing a drug to market. The chart below is one of the clearest visualizations I’ve seen of this imbalance. Discovery-to-lead—the dark-blue sliver on the left—covers identifying promising targets and optimizing them into drugs worth advancing. Everything after it—confirming the lab tests and cells used to evaluate a candidate are accurate and reliable (assay and cell line validation), then manufacturing, animal toxicology, and three phases of human trials—occupies nearly 99 percent of the remaining bar.
That disparity suggests two things.
First, if AI models commoditize, many standalone AI-discovery companies may be worth little more than a ChatGPT enterprise subscription and the pitch deck they neatly wrapped around it. A full-stack biotech company can use its own AI models to prioritize compounds and conduct the full array of testing in its own laboratory. AI becomes a tool inside the business rather than the business itself.
Second, computational abundance makes skilled practitioners more valuable. While it’s easy for a model to propose a molecule, it takes medicinal chemists and toxicologists to decide whether the evidence is strong enough—and the risks tolerable enough—to justify putting it into a living person. The judgment and, in many ways, the taste, to know which drug is worth backing is much harder to replicate and scale, and is why expertise and conviction remain a huge bottleneck in drug development.
There will be exceptions, of course. An AI company can capture the upside if it owns the lab work and the drugs, but only by accepting the cost and risk of failure that every other biotechnology company bears when bringing a drug to market. At that point, it becomes a pharmacology company.
The future may look less like software eating pharma than pharma eating software.—Ashwin Sharma
That’s all for this week! Be sure to follow Every on X at @every and on LinkedIn.
Everyone’s a builder now. Every All Access gets you the full membership plus the Builder Pack—$7,000+ in credits for the tools we build with.















Comments