We use analytics and advertising tools by default. You can update this anytime.
We revisit Anthropic’s newest model, Figma responds to the 'SaaSpocalypse,' and how Every’s senior designer wrangles multiple image generators
Today, we update our Opus 4.8 Vibe Check with a Pulse Check featuring perspectives from more team members, Dan Shipper sits down with Figma’s Matt Colyer to unpack why AI hasn’t killed professional design services, and Every senior designer Daniel Rodrigues shares the two-tool AI workflow he uses to get precise, visually stunning results.
Was this newsletter forwarded to you? Sign up to get it in your inbox.
In a new episode of our podcast, AI & I, Dan talks with Matt Colyer, Figma’s director of product management for developers, about the limits of chat-based AI agents for design and why the rise of vibe-coded everything is, despite what you might have heard, a boon for the company.
Watch on X or YouTube, or listen on Spotify or Apple Podcasts. (You can also read the transcript.)
Here are the highlights:
Miss an episode? Catch up on Dan’s recent conversations with LinkedIn cofounder Reid Hoffman; the team that built Claude Code, Cat Wu and Boris Cherny; Vercel cofounder Guillermo Rauch; podcaster Dwarkesh Patel; and others, and learn how they use AI to think, create, and relate.
You’re probably used to old product specs. You write acceptance criteria, engineers build according to it, and QA verifies that it shipped correctly. But AI doesn’t do that—it gives different results every time. Braintrust just published “Evals Are the New PRD”—the argument is that, for AI products, evals replace the spec, the acceptance criteria, and the roadmap all at once. While a PRD gathers dust in a Google Doc, an eval suite runs on every commit. The piece walks through a four-stage flywheel: Observe, analyze, evaluate, improve. It’s based on how teams at Stripe, Zapier, and Vercel actually ship quality AI. Read it now.
Five days ago, we called Anthropic’s Claude Opus 4.8 the best Claude model yet for writing and serious engineering, and said we’d switch to it from GPT-5.5 if the Claude app ever caught up to Codex. After a work week of more testing, we’re still an Opus 4.8 admiration society, although the results are a bit more mixed as people from different disciplines have had a chance to weigh in.
Here’s what more of the Every team has to say about when to use the model and when to steer clear.
Arielle Shipper, Every’s new head of operations, has spent the last few weeks on a discovery tour. She used Opus 4.8 to redo an HTML site showing a summary of her findings, after building the original with Opus 4.7. She noticed meaningful improvements: 4.8 distinguished between two similarly named pages in Notion without the explicit guidance 4.7 had required, and suggested highlighting a count of how many times specific topics came up in her conversations with the team. Her summary: “It seems really detail-oriented in a way I appreciate.”
Austin spent the weekend using Opus 4.8 on an essay with Monologue, our speech-to-text tool, and our writing app, Spiral. For that job, he wrote that Opus 4.8 “is the best model available,” a step up from Opus 4.7 and “materially better than GPT-5.5.” But he doesn’t expect it to change his daily behavior. GPT-5.5 is “pretty good” at the same kind of creative partnership, he said, and keeping his work in Codex matters more than the modest quality improvement: “I don’t see myself reaching for Claude models much without a materially better desktop app experience, or such a dramatic leap in model quality that the harness matters less.”
Nityesh tested Opus 4.8 inside the AI employees he is building for Every—Claudie for consulting, Andy for the editorial team. He reported that the model recalls the right memory at the right time, stays useful in longer threads, and lets him use more of its 1-million-token context window, the amount of material it can handle in one conversation. But Anthropic really won his heart with Dynamic Workflows, the workflow-automation feature released alongside Opus 4.8. Combined with the new model, Nityesh says it feels like “a major power-up.”
Anthropic says Opus 4.8 is more honest and better at flagging risks. But Lee saw the negative side of that instinct during a daily planning run he’d repeated for months where Claude used his calendar, Slack, and notes to create a plan for his day. One morning, the plan cited events, messages, and files Lee couldn’t find in those sources. When he asked Claude what had happened, it claimed a prompt-injection attack had supplied fake information. When Lee challenged it, Claude said it had invented that story to explain its own bad output, mistaking a planning file Lee had moved for evidence of interference. The exchange left him reluctant to trust the model’s explanations for its own behavior.
Andrey is “very positive” about Opus 4.8 for coding and wrote that he likes it much more than GPT-5.5. For his use cases, it feels “more stable, reliable, and just less dumb.” His reservations are about the experience around the model, not its coding quality: GPT-5.5 is faster, and Codex gives it the better desktop-app harness.
Become a paid subscriber to Every to unlock this piece and learn about:
You’re probably used to old product specs. You write acceptance criteria, engineers build according to it, and QA verifies that it shipped correctly. But AI doesn’t do that—it gives different results every time. Braintrust just published “Evals Are the New PRD”—the argument is that, for AI products, evals replace the spec, the acceptance criteria, and the roadmap all at once. While a PRD gathers dust in a Google Doc, an eval suite runs on every commit. The piece walks through a four-stage flywheel: Observe, analyze, evaluate, improve. It’s based on how teams at Stripe, Zapier, and Vercel actually ship quality AI. Read it now.
Join 100,000+ leaders, builders, and innovators

Already have an account? Sign in.
Daily insights from AI pioneers + early access to powerful AI tools
Comments