We use analytics and advertising tools by default. You can update this anytime.

Plus a workflow for managing agents while doing the dishes, why AI needs to be more social, and who to follow on X for design inspiration
If you’ve been following our coverage—or CEO Dan Shipper on X—you know the Every team is going all in on voice. On Friday, August 7, we’re hosting a camp for paid subscribers about all the tangible ways we’re using ChatGPT voice mode to get stuff done away from the keyboard.
Was this newsletter forwarded to you? Sign up to get it in your inbox.
Last week, voice mode was blowing up Every’s Slack.
Powered by GPT-Live, OpenAI’s new voice model, the feature lets you have natural conversations with ChatGPT, complete with interruptions, follow-up questions, and redirections. Within the ChatGPT desktop app, voice mode can find the right task or thread based on spoken context, kick off new threads, check on existing work, and send more complex tasks to GPT-5.5.
After Dan took to X to evangelize voice mode’s powers for writing and revising an essay, the team put the feature through its paces. We used it to fix user-reported bugs, draft article outlines, do meal prep, draw connections between what we were reading and what we were building, book airline tickets, and orchestrate agents while cooking.
What works: There’s a lot to love about voice mode, which allows you to get work done without having to sit at a keyboard.
One of its biggest strengths is that it lets you read and ask questions aloud or connect what you’re reading to another file or project. Engineer Lee Knowlton uploaded a PDF of Designing Data-Intensive Applications, a book about building large-scale data systems, while voice mode had access to the codebase he was working on. As a result, he could keep reading while asking questions aloud, exploring unfamiliar ideas, and connecting the book’s insights to his own code. The result was a more fluid way to learn.
“Shifting from text to voice is different from shifting from text to text for me,” he says. “Reading something and then having a conversation, or asking a quick question, is different from typing something and then having to parse more text.”
What could be better: During a walk, COO Brandon Gell found that the mobile app’s voice mode could read a thread’s visible history but didn’t have access to important context outside the thread. Currently, voice can control local Codex work through Remote connections—but only while the host computer is awake, online, and running the desktop app. Without that connection, voice can’t access the host’s Codex projects, files, or tools.
Further complicating matters, the mobile app also has an “ordinary voice mode,” which can use the current cloud conversation but not the local context available through Remote. Are you confused? We’re confused.
The model’s ability to distinguish between speech intended for it and ambient conversation was also inconsistent. Lee found it good at filtering out exchanges with his wife, while engineer Tyler Nishida had the exact opposite experience. And the lag time can make it hard to use as a writing or editing partner (I found voice mode impressive but functionally too laggy to help me write this piece, for example). Finally, although GPT‑Live can delegate complex tasks to a frontier model in the background, some responses still felt shallow compared with responses from a text chat set to GPT-5.6 Sol.
Final verdict: Voice mode is, as Dan puts it, “a whole new world”—one in which you can direct agents away from a computer. But there are kinks to work out.
“It’s both not quite there yet and obviously the future,” Lee says. “A week ago, I couldn’t imagine a version of this that was really good, and now I can.”
Customer service volume doesn’t scale with headcount—ElevenAgents does. The platform deploys voice and chat agents that handle billing questions, support tickets, and outreach around the clock across 70+ languages, integrating with Salesforce, Zendesk, and the tools your team already runs. Faster resolutions mean fewer abandoned carts and higher LTV.
Sarah Tavel has spent her career studying consumer technology cycles, and she thinks she’s spotted the next one. A former Pinterest product manager and current Benchmark partner, she’s betting the next wave of AI products won’t just be smarter, they’ll be social. On this week’s AI & I, we’re revisiting our April 2025 conversation with Sarah, who argues that even power users are still using AI products like ChatGPT in a rudimentary way. But the gap isn’t the models: It’s that nobody has built a way for users to learn from each other. Sarah thinks whoever captures and shares that knowledge will create the next big product.
Watch on X or YouTube, or listen on Spotify or Apple Podcasts. You can also read the transcript.
Here are the highlights:
This episode is a must listen for anyone who wants to understand why the biggest AI product hasn’t been built yet—and what it might take to build it.
Miss an episode? Catch up on Dan’s recent conversations with Anthropic head of product Mike Krieger; the team that built Claude Code, Cat Wu and Boris Cherny; the team that built Codex, Thibault Sottiaux and Andrew Ambrosino; Vercel cofounder Guillermo Rauch; podcaster Dwarkesh Patel; and others to learn how they use AI to think, create, and relate.—Miriam Partington
Lee’s workflow shows how voice mode changes the way we interact with agents...
Become a paid subscriber to Every to unlock this piece and learn about:
Customer service volume doesn’t scale with headcount—ElevenAgents does. The platform deploys voice and chat agents that handle billing questions, support tickets, and outreach around the clock across 70+ languages, integrating with Salesforce, Zendesk, and the tools your team already runs. Faster resolutions mean fewer abandoned carts and higher LTV.
Join 100,000+ leaders, builders, and innovators

Already have an account? Sign in.
Daily insights from AI pioneers + early access to powerful AI tools