
From Doing to Tending
Plus: A mini-Vibe Check of Grok 4.5, and why AI scribes could dull doctors' judgment
Was this newsletter forwarded to you? Sign up to get it in your inbox.
Mini-Vibe Check: Grok 4.5 is fast, cheap, and finally useful
Almost a year ago, we vibe-checked Grok 4 and found a model that performed well on benchmarks but was not useful enough for our engineers to use every day. Our April Vibe Check of Cursor 3.0 found a promising but unfinished agent-orchestration product, but that verdict aged quickly: Kieran Klaassen was using Cursor daily by the end of the month, and by June Composer 2.5 was his main model for final polish. Then, in June, SpaceX signed an agreement to acquire Cursor, and a new candidate for the AI frontier space race was born.
Grok 4.5 is the first result of that collaboration. Cursor says it jointly trained the model with SpaceXAI using data from interactions with codebases and software tools.
When our team ran Grok 4.5 through the evals we use internally, the consensus was that it is an Opus-level model. Mike Taylor’s latest benchmark put it slightly above Claude Opus 4.8: Grok followed every step and returned a complete, polished result, while Opus stopped early or skipped parts of the assignment.
Create a free account, or log in.
Every members live and work at the edge of AI. Join now.
By continuing, you agree to the Terms of Sale, Terms of Service, and Privacy Policy.
Enjoy unlimited access to all of Every.
See subscription options