
The Ur-model Cometh
Plus: What’s next for Every consulting
Hello, and happy Sunday! This past week was like Christmas in February. Anthropic and OpenAI both dropped significantly improved models that moved us closer to peak general-purpose AI, and it was all-hands here at Every to share what our advance testing revealed about each of them. The result was Vibe Checks on both Opus 4.6 and Codex 5.3, the inevitable head-to-head showdown, and a livestream featuring Sam Altman himself.—Kate Lee
Was this newsletter forwarded to you? Sign up to get it in your inbox.
Knowledge base
“GPT-5.3 Codex vs. Opus 4.6: The Great Convergence” by Dan Shipper/Vibe Check: Opus picked up Codex’s precision. Codex gained Opus’s warmth and willingness. Every CEO Dan Shipper and the team tested both extensively, and the verdict is that these models are—in a good way—beginning to resemble each other. Most of the Every team is now using both. Read this for the full head-to-head breakdown, including which one has the higher ceiling and which delivers steadier, faster autonomous execution.
“Vibe Check: Opus 4.6—The Best Coding Model We’ve Tested (With Some Maddening Habits)” by Dan Shipper and Katie Parrott/Vibe Check: In 15 minutes, Opus 4.6 solved a Monologue iOS problem that stumped both Codex and Opus 4.5—researching competitors and open-source repos to find the perfect solution. Put simply: It’s extremely smart. As proof, it set the high score on Cora general manager’s Kieran Klaassen’s LFG benchmark. Some trade-offs exist, though—it’s slower and occasionally confabulates, and the team preferred Opus 4.5’s prose in blind tests. But for vibe coders? Switch now. Read this to learn why.
Create a free account, or log in.
Every members live and work at the edge of AI. Join now.
By continuing, you agree to the Terms of Sale, Terms of Service, and Privacy Policy.
Enjoy unlimited access to all of Every.
See subscription options