
Vibe Check
GPT-5
Our hands-on review of OpenAI's newest model based on weeks of testing
The Verdicts
Legend
GPT-5 in ChatGPT
OpenAI can accomplish this because GPT-5 isn't one model, it's a system of models. In ChatGPT, an "auto-switcher" determines the intent of your query, and decides whether to route it to the chat (for easy queries) or reasoning (for more challenging questions) version of the model.
For questions that it can answer off the top of its head—like, "What is the definition of AGI"—it will respond lightning-fast. For more complex queries—like, "Code me a beautiful new social network"—it will think for a while or do web research before returning a more considered answer.
In ChatGPT, GPT-5 returns comprehensive and readable answers with logical subsections, judicious use of white space, and bolded text to help you find what you need quickly. (It will remind you of answers you get from o3, but without the obsession with tables.)
However, my results vary in the non-reasoning version of the model. Sometimes it's great, but it often hallucinates on questions that should have been routed to the reasoning model. For example, if I take a picture of a passage in a novel and ask it to explain what's happening, GPT-5 will sometimes confidently make things up. If I ask it to "think longer," it will deliver an accurate answer.
In ChatGPT's collaborative workspace Canvas, GPT-5 can quickly one-shot front-end apps, making it the introduction of vibe coding to millions of people who would never pay for a Claude subscription or try coding app Loveable. But it's not yet a game changer if you're already a vibe coding veteran: Its work is about on par with Opus 4.1 in Claude Artifacts. Canvas also has a number of quirks that make it hard to work with; for example, it is limited to fewer than 1,000 lines of code.
The bottom line: This will be the first time most of the world has ever used a reasoning model instead of a simple chat model. GPT-5 is available for free, for everyone, today—that's a big deal.
GPT-5 in the API
For comparison, Google's cheap and fast model, Gemini 2.5 Flash, costs $0.30 per million input tokens. As a result, it's one of our favorites to use at Every. But OpenAI now has a direct answer: GPT-5-mini, which clocks in at $0.25 per million input tokens, undercutting Flash.
At the flagship tier, GPT-5's Standard pricing is $1.25 per 1 million input tokens—exactly matching Google's Gemini 2.5 Pro. If you're already paying for Pro-tier Gemini, you can switch to GPT-5 without changing your unit economics.
On the Anthropic side, the comparison is almost comical. Claude 4 Opus is pegged at $15 per million input tokens. GPT-5 Standard is $1.25 per million. That's 12 times cheaper.
Even if Opus 4.1 is better for some use cases than GPT-5, it's difficult to be 12 times better. GPT-5 forces a hard look at the math.
GPT-5 for agentic engineering
Don't get me wrong: GPT-5 is a very good programmer. It's incredibly useful as a pair programmer, especially in AI-powered integrated development environments (IDEs) like Cursor. It's great for engineers from traditional backgrounds who want an AI to help collaborate on code. And it excels at research and debugging complex issues.
But the discipline of programming has fundamentally changed this summer. The benchmarks don't show it, but if you know how to YOLO four agents at once in Claude Code, GPT-5 feels like a step backward. That's partially because of the model's current personality: It's more cautious than Opus 4.1 and isn't as comfortable working independently for long periods in our testing. But it's also due to the app you use to interact with it: Both Cursor and OpenAI's command line interface tool Codex CLI are not on the same level as Claude Code. Both were built for programmer-AI pair programming, not true delegation.
I bet this will change. The model is extremely smart, just not yet built for this use case. But for now, OpenAI seems to have missed the paradigm shift in programming caused by Claude Code over the last two months.
Table of Contents
The reach test


For day-to-day tasks
It's a daily driver in ChatGPT. We never use the model picker anymore and almost never need to go back to older models. It's extremely fast for day-to-day queries and gives comprehensive answers to questions that require research. It also disagrees more frequently instead of hallucinating.
For pair programming
Use it to fix a specific bug or build a new feature step-by-step with a helpful companion in Cursor. It's great at researching and understanding large codebases. It's extremely detail-oriented and fast.
For writing
GPT-5 has a good voice—nuanced and expressive. It's less likely to output obvious AI idioms, so it's the first thing we turn to when we have a sentence we need to polish or a paragraph we need to draft. We sometimes return to GPT 4.5 for questions that require more thought.
For agentic engineering
GPT-5 in Codex and Cursor is too cautious to be a good agentic programmer. It stops too often, and its output on front-end and back-end tasks is lower quality than that of Opus 4 and 4.1 in Claude Code. On big tasks, it gets lost in the details, and its output tends to be too verbose to read easily.
For editing
GPT-5 cannot determine whether writing is good. We have benchmarks (below) to test AI's ability to judge writing, and GPT-5 consistently fails on tasks that Opus 4 passes.
The team roundtable
GPT-5 is a Sonnet 3.5 killer, not a leap into the future.
I was working on a feature inside of Spiral that used an open-source framework that couldn't quite do what I needed. I used GPT-5 to merge code from another open-source framework into the one I was using. It didn't do it in one shot, with just one example, but something about the process felt viscerally collaborative. I felt this growing sense of confidence that we were getting there together.
GPT-5 has become my go-to for discrete, well-defined coding tasks. I still use Claude Code for longer-running, more agentic work like code reviews, but if I'm blocked or too lazy to fully think something through, working with GPT-5 gets me where I need to go.
For developers, at $1.25 input and $10 output per 1 million tokens, GPT-5's sweet spot is when I've crafted a solid prompt and need to process lots of information into a concise and high-quality output, such as documentation or a new coding function. It's dramatically cheaper than Opus but pricier on outputs than o4-mini, so you're paying for steerability—its ability to adhere to your prompts—not raw reasoning (where o3 may still win). GPT-5-mini might be the real surprise, though, undercutting Gemini's Flash on price with similar performance, assuming it can match the speed.
I've been way happier overall with GPT-5 over Opus as a first draft writer.
For the challenging research task of finding duplicate files on a Mac, it gave me the most technically rigorous writeup I've ever seen from an AI. It was like talking to a 140-IQ systems architect who's already built the network three times and learned from each failure.
For pure implementation, I'll still reach for Claude. But when I need deep context, tradeoff analysis, and "why" answers that change the way I build features, GPT-5 is unmatched. I wouldn't go back to Claude for research.
The benchmarks
Good writing
GPT-5 produces inconsistent results, sometimes passing and other times failing the same piece of writing. It was inconsistent enough for Danny to not fully trust its evaluations, especially compared to Claude Opus, which reliably gives the same results every time.
We ran it on a series of writing samples from tweets to essays, and it consistently returned "false," judging the writing engaging when it wasn't:
Danny also ran a blind "taste test" between Opus 4 and GPT-5. He gave both models the same set of prompts and asked the Every team on Discord to vote on the outputs, without revealing which model wrote what. Opus 4 (and later, 4.1, which Anthropic dropped earlier this week) came out on top.

One-shot a game
These are screenshots from the game that GPT-5 made:


For comparison, this is the output of Opus 4.1 in one shot:

AI Diplomacy
Using an early GPT-5 variant that does minimal reasoning, it ranked near the bottom with the baseline prompt but jumped to second place when told to be aggressive, cutting "hold" moves from 49 percent to 9 percent—even better steerability (read: prompt following) than o3. The public GPT-5 is a slower, more reasoning-heavy version that performed worse in this benchmark.
This chart shows the results from 20 games where various models acting as France faced weaker opponents. The red bars represent optimized prompts, the gray basic prompts, and the yellow the average (lines show range). o3 leads overall, but GPT-5 variants compete well with quality prompts—note GPT-5-mini matching Flash in the orange outline. The red/gray gap demonstrates how much prompt engineering matters for performance.
Alex's takeaway: It's great for consumers and very steerable, so great prompts give great results. But it's not a frontier-pushing release, at least on this benchmark.

Impossible puzzle
He's tried to solve it with every new model. GPT-5 solved it in 1 minute and 10 seconds. The only other models that have been able to do so are o3 and o3 Pro, which took 8 minutes and 19 minutes, respectively.

One-shot a music production app


Pelican on a bicycle
Someone call Microsoft: OpenAI has achieved AGI internally.

'Thup'

