Skip to content
Why Evals Are So Hot Right Now
Midjourney/Every illustration.

Why Evals Are So Hot Right Now

Plus: Where we stand on Grok 4.7, a workflow for going viral on social media, and why every company needs a Slack agent

Sep 24, 2026 · 9 min readUpdated Oct 1, 2026

Comments1

In today’s Context Window, we review Grok 4.7 (“a step backward”?) and explain why evals are suddenly everywhere. Elsewhere, head of social media Becky Isjwara breaks down her workflow for turning Every articles into viral short videos, head of platform Willie Williams argues that Slack agents are the new websites, and we share a fresh crop of Thesis Statements, including entries from Cmpnd’s Drew Breunig, Notion’s Geoffrey Litt, and MIT’s Daniela Rus, about what happens after automation.

Was this newsletter forwarded to you? Sign up to get it in your inbox.


Mini-Vibe Check

Grok 4.7 is uneven

Grok 4.7 launched publicly on Monday. And the vibes? Well, they are evolving.

In the team’s early tests, which began before the public release, the model was practically unusable. It left tasks unfinished, ignored instructions, and merged code without permission.

Every has numerous Grok 4.6 power users—notably Cora general manager Kieran Klaassen and designer Tyler Nishida—who like it because it’s cheap, fast, and reliable enough to be a workhorse. (Tyler estimates Grok accounted for 97 percent of his tracked code events last month.)

Grok 4.7 initially felt like a major regression. When he ran the model against his personal benchmark last week, head of evals Mike Taylor reported that it scored 42 percent compared with Grok 4.6’s 84 percent. “Maybe they rushed it out, or they have something better coming down the pipe,” he says. “This seems like a step backward.”

Case closed. Or was it?

It was not, at least not according to Tyler, who noticed that the model’s performance improved leading up to its public launch. In particular, GrokBot, the company’s AI assistant, got better after it began using Grok 4.7. Where he once sent a string of all-caps corrections—“Stop, stop, stop. That’s not what I wanted!”—he now sends simple one-line requests. “I feel like I’m micromanaging it a lot less,” he says.

By last Friday, Grok 4.7 was working well for him. As of this writing, it’s his daily driver. He even prefers it slightly to Grok 4.6—it retains what drew him to Grok, a cheap, steerable model that lets him keep working without constantly hitting usage caps.

His endorsement has limits: He still finds Grok 4.7 frustrating when it coordinates other agents and prefers using it in a single chat. And Elon Musk—who, in a (since-deleted?) tweet, declared Grok 4.7 would “exceed all current models”—had led him to expect more than an incremental improvement. “I wouldn’t have been this excited for it, and I feel just kinda let down,” Tyler says.

Grok 4.7 did not live up to Elon’s hype. (Screenshot courtesy of Tyler Nishida.)
Grok 4.7 did not live up to Elon’s hype. (Screenshot courtesy of Tyler Nishida.)


Tyler’s teammates haven’t all come around. Kieran’s assessment remains negative. “Compared to Opus 5.5, this model is not even close,” he says. “It’s not a model I would even consider using.”

Staff writer Katie Parrott’s writing tests were also discouraging. She dislikes Grok 4.7’s editorial judgment. Asked to rewrite the introduction to Dan Shipper’s “After Automation,” it produced the line “If the model can do the task, headcount is sentiment,” an example of the model’s “big Elon Musk energy,” Katie says.

She found its writing style equally poor. In that same “After Automation” rewrite, its opening paragraph consisted entirely of short declarative sentences.

A Grok 4.7 editorial massacre. (Screenshot courtesy of Laura Entis.)
A Grok 4.7 editorial massacre. (Screenshot courtesy of Laura Entis.)


“It has no cadence or rhythm and no empathy or theory of mind for the reader,” she says.

Our verdict: Grok 4.7 has won back at least one of its biggest users at Every. If you liked 4.6 for everyday coding, Tyler’s experience gives you a reason to try the upgrade. If you’re using AI to write, however, Katie’s tests suggest you leave the model alone.


Signal

Evals go mainstream

What happened: Evals are having a moment. Of the 25 product management openings tech influencer Lenny Rachitsky shared last week, he says nearly half asked for experience writing evals, or tests of how well an AI system performs a task.

At app monitoring company Sentry, proposed code changes triggered roughly 1,800 eval runs in the past month, engineer Ryan Brooks reported—up from near zero in May. These checks can catch problems before an update ships, such as revised instructions that cause an agent to struggle with tasks it previously handled well.

What it means: Choosing the most intelligent model used to be a reasonable default, says senior applied AI engineer Nityesh Agarwal. But frontier models have become so capable that most knowledge workers don’t need their full intelligence for everyday work. Several models can already handle those tasks, making speed, price, and alignment with the user’s taste more important selection criteria.

Nityesh is helping everyone at Every build personal benchmarks based on existing tasks and individual standards for good work. As model releases accelerate, these benchmarks let us quickly test how well a model handles different parts of our jobs. “If I have my own personal benchmark, I’ll know whether I need to pay attention to a model release or not,” Nityesh says.

Personal benchmarks measure how good AI is at your specific job. (Screenshot courtesy of Laura Entis.)
Personal benchmarks measure how good AI is at your specific job. (Screenshot courtesy of Laura Entis.)


Why it matters: Teams building with AI increasingly need to judge its quality for themselves by defining success and checking whether model updates help or hurt their work.

Evals help teams catch failures before customers encounter them and test whether a faster, cheaper model performs well enough for their product. For individuals, evals identify which model works best for a specific task based on its capability, speed, and cost. As models improve, “you don’t necessarily need your daily driver to be the most intelligent one,” Nityesh says.


Steal this workflow

Turn an article into a short video

Every’s head of social media Becky Isjwara has found viral success turning articles into short animated videos. For Katie’s Compound Writing guide—which explains how to write with AI to improve your craft over time—she made an 18-second clip that animated the cover’s flying papers and illustrated the writing process as a loop.

Becky Isjwara's animation built from the Compound Writing cover image. (Video courtesy Becky Isjwara.)

To make these videos, Becky uses Hyperframes with Fable in the Claude desktop app. Hyperframes lets Claude turn text, images, and animation into a finished video.

Here’s her process:

Step 1. Decide how to represent the article visually. Give Claude the article URL and ask it to suggest an opening hook and a sequence of scenes. Read the piece yourself, then check whether Claude’s proposed scenes show the article’s most interesting ideas.

Step 2. Plan each scene. Ask Claude to create a storyboard with one still image for each scene. Have it use colors from the article’s cover image and include the text that will appear onscreen. Before moving to animation, check that the story is easy to follow and the visuals convey the intended ideas.

Step 3. Make the video. Once you’ve approved the storyboard, ask Claude to turn those scenes into a video with Hyperframes, adding motion and transitions. If you want to animate the cover illustration, upload it to Flora, a workspace for creating images and videos with AI. In Flora, select MiniMax, an AI video generation model, and describe how you want the illustration to move. Then give Claude the resulting clip and ask it to include it in your storyboard.

Step 4. Watch and refine. Watch the video in Claude’s built-in preview, pausing to tell Claude what to change—for example, “Keep this scene on screen longer so I can read the text” or “The last letter is cut off. Adjust the layout so the whole word is visible.” Have Claude apply your requests, then review the video to confirm the changes and check for any new problems.

Try it this week: Choose an article and a cover image, and use these steps to make a short video. Once you’re happy with the result, ask Claude to save the workflow and your corrections as a reusable skill. Instruct it to seek your approval at each step, so you can make a video for a new article without repeating these instructions.


Thesis Statements

Last month, we launched Thesis Statements, a collection of specific, contestable claims from builders and thinkers about the future of great human work with AI.

Here are five more predictions from people at the frontier:

If you want to help decide what matters in the future of AI and human work, think creatively, and build what comes next, join us at our inaugural Thesis: 2027 conference on November 5, 2026.


Jagged frontier

Slack agents are the new websites

In the internet’s early years, any company moving online first needed a website. Websites started as informational front doors—a place to say what you did, list a phone number, and maybe show some photos. Then the website became the product—you could order directly from walmart.com, rather than just viewing a list of store locations. Websites turned into web apps, and then we forgot that they were ever different.

The same thing is about to happen with agents.

Applications are evolving beyond websites into agents on conversational platforms like Slack, Teams, Google Chat, WhatsApp, Telegram, and whatever comes next. These surfaces already hold much of the context agents need. Putting agents there lets them use that context directly rather than move it elsewhere.

Take Every. To read one of our articles, you need to visit our site, search for it, and copy anything you want to use into another tool. An Every agent that lives in Slack (ed: like the one we’ll soon be launching) could access our archive, bring you relevant pieces, and discuss them with you where you’re already working.

Too few companies recognize the opportunity to deploy their own branded agents where their customers communicate. Websites went from optional to essential for businesses in a decade. Slack agents could make that leap even faster.

Agent frameworks are making these tools cheaper to build. As more companies launch agents, competitors will feel pressure to follow.

The next front door to your business is an @.—Willie Williams


Laura Entis is a staff writer at Every. To read more essays like this, subscribe to Every, and follow us on X at @every and on LinkedIn.

Everyone’s a builder now. Every All Access gets you the full membership plus the Builder Pack—$9,000+ in credits for the tools we build with.

Related Essays

Comments

You need to login before you can comment. Don't have an account? Sign up!
Every

A paid Every subscriber gifted you this article. Every is a bundle of writing and tools for making sense of AI, work, and what comes next.

Subscribe to the Every bundle

We use analytics and advertising tools by default. You can update this anytime.