We use analytics and advertising tools by default. You can update this anytime.

Plus: Why intent beats volume, a new AI productivity status metric, and a tool for squeezing the most out of cheaper models
For a brief period, AI was an affordable novelty. Every use case felt like magic, and even the silliest task was worth a try when frontier labs were subsidizing compute costs to get consumers hooked. Power users proved their status by maxing out their token consumption.
Now, AI is ubiquitous—even required in many workplaces—and easy to use for anything from code to text to visuals. But this flood of production has a price: in cash, as powerful new models grow more token-hungry and the labs roll back those generous subsidies, but also in the time and effort required to make sense of the results. (If you’ve ever tried to debug an AI-generated codebase, edit AI-generated text, or decipher the meaning behind an AI-generated email, you know how labor-intensive it can be to wade through poor-quality LLM outputs.)
Focus has evolved accordingly from how much you’re using AI to how you’re using it—and what you can show for it. Does the benefit justify the significant cost?
Today’s Context Window explores various answers and solutions to that question. First up, author and technologist Craig Mod explains why cheap software creation has made him more protective of his writing time; Monologue general manager Naveen Naidu shares a new efficiency metric he heard making the rounds in San Francisco; senior applied AI engineer Nityesh Agarwal shows us how he audits agents for wasted tokens; and Spiral general manager Marcus Moretti explains how OpenRouter helps him manage a 12-plus-model stack.
Was this newsletter forwarded to you? Sign up to get it in your inbox.
The entire economy is up for grabs in a post-AI world. What will you build?
Elbow Grease is a new accelerator from Gutter Capital in NYC. This is not a finishing school for fundraising—it’s where you build a business with people who’ve done it before. The program offers a $300,000 initial investment, weekly coaching from Gutter partners Dan Teran and James Gettinger, and 1:1 mentorship from a Series B+ or exited founder. Come work among the Gutter portfolio and learn from industry veterans Scott Belsky, Gokul Rajaram, and Hunter Walk, to name a few. Apply by July 31!
AI is powerful. According to CEO Dan Shipper, it’s also a “slot machine.” How, then, can you use the technology to create stuff that matters while avoiding the hunt for the next dopamine hit?
To help answer that question, Dan had author and technology enthusiast Craig Mod on the show to discuss how to be ruthless about preserving his time.
Watch on X or YouTube, or listen on Spotify or Apple Podcasts. You can also read the transcript.
Somewhat counterintuitively, the ease with which software can be made now has reaffirmed Mod’s commitment to writing. “There are plenty of people playing around with this stuff,” he says. “But there aren’t that many people who are going to think about or write the weird books I feel drawn to write, and as a human, that feels like the valuable thing for me to put my effort into.”
While he’ll use AI for research and fact-checking, he still writes every word himself. Outsourcing that process to an LLM would defeat the purpose when “being in the mess of writing” is the point—and the way to get the results he’s looking for.
Miss an episode? Catch up on Dan’s recent conversations with LinkedIn cofounder Reid Hoffman; the team that built Claude Code, Cat Wu and Boris Cherny; Vercel cofounder Guillermo Rauch; podcaster Dwarkesh Patel; and others, and learn how they use AI to think, create, and relate.
While on the ground in San Francisco for Apple’s Worldwide Developers Conference last month, Monologue general manager Naveen Naidu noticed a new metric for measuring enterprise productivity making the tech-circle rounds: revenue per million tokens.
A purported measure of a company’s efficiency, revenue per million tokens is a potential successor to revenue per employee, a rough estimate for how much money each worker at a company generates. (AI-native companies tend to score significantly higher on this scorecard than their traditional SaaS counterparts.)
By explicitly tying ROI to how efficiently a company makes money with AI, revenue per million tokens acknowledges that engineering has become cheap while the cost of compute is more expensive than ever. Or more simply: “If you say, ‘I wrote a million lines of code,’ did it actually increase your revenue or not?” Naveen says.
I’ve started to think of senior applied engineer Nityesh Agarwal as Every’s resident AI seer, particularly when it comes to Anthropic. Three to six months before the AI lab released a strategy to handle the problem of agent orchestration and team-based agent delegation, Nityesh had built his own version of a solution. He’s been sharp and early on understanding the architecture that allows an agent to access the context it needs to get work done on behalf of an entire team. (If you don’t believe me, read his piece on how Claude Code is a more reliable alternative to OpenClaw, written before pretty much anyone else saw its potential.)
So where is Nityesh focused now that Anthropic has essentially commoditized the model of shared, Slack-based team agents (like the ones he built for Every’s consulting and editorial teams)?
“Making the agent more token efficient,” he responds without hesitation.
Up until now, you could throw money at AI experiments because the cost was largely covered by a subscription model that subsidized heavy use under a set monthly price. Those days are numbered, Nityesh says. “We are in a compute-constrained world where we will not have enough data centers to run tomorrow’s AI.” Expect usage costs to go up, particularly for individuals as AI companies prioritize big-budget enterprise customers.
At the same time, cheaper models—many of them open-sourced—have grown more capable, operating, by Nityesh’s estimation, roughly seven months behind the frontier.
Here’s how Nityesh is thinking about token efficiency:
Not everything done inside the AI writing assistant requires Opus 4.8-grade intelligence—or cost. An LLM gateway, OpenRouter helps Spiral general manager Marcus Moretti use the right-sized model for the right task.
Spiral currently uses 12 different models, including Sonnet 4.6 for most prose, Gemini 2.5 Flash for a top-edit that removes AI tells, and a smaller, lower-cost OpenAI model for summaries of files. Staying on top of each provider’s API format, credentials, account, and billing would be complicated and time-consuming.
OpenRouter takes care of that management layer for him—once integrated, the service gives Spiral a standardized way to send requests to many different models. Marcus regularly checks out OpenRouter’s LLM Leaderboard, which ranks models by weekly token usage across the platform and, of late, has been dominated by cheaper, open-source options rather than expensive frontier LLMs from OpenAI and Anthropic.
Another big plus—and the original reason Marcus turned to OpenRouter—is reliability: It makes models available through multiple providers. If Anthropic runs into an issue, for example, OpenRouter can tap another provider, such as Google or AWS, so users don’t experience any interruption in service.
It’s causing tech worker sentiment to split into two. (Lenny’s Newsletter) Entrepreneur and Substacker Lenny Rachitsky just published his second annual tech worker survey, which revealed “a tale of two workforces.” Roughly half of respondents experience AI as an amplifying, energizing force, while the rest feel shaken and destabilized by it. There are meaty insights in here, including a surge in people reporting “significant” burnout, how even tech workers optimistic about their professional trajectories wouldn’t recommend their career path to newcomers, and mixed emotions about the overall impact of AI on work. “The defining feeling about AI is ambivalence,” Lenny writes.
AI CEOs want you to know they were maybe wrong about mass unemployment. (Wall Street Journal) After years sounding the alarm on the economic implications of their frontier models, AI CEOs are striking a more optimistic tone. “Our industry underestimated how much we’re going to be able to keep people at the center of everything,” per OpenAI CEO Sam Altman. Meanwhile, entry-level job doomer (and Anthropic CEO) Dario Amodei recently acknowledged AI-motivated layoffs could be avoided should companies apply “creativity” to achieving more with the same resources. What’s behind the about-face? Per the WSJ, it could be a better understanding of AI’s role in the workplace, a PR tactic in response to growing public backlash, or a blend of both.
The supporting data is a giant question mark. (New York Times) No one is arguing AI isn’t impacting the labor market. What’s up for debate is how: Depending on your data source, the technology is either destroying or creating jobs, while contributing to or helping solve inflation. Big picture, there’s a limit to what current, inherently incomplete metrics can tell us about how AI will transform work. “What the data can almost never tell us is where we’re going to be in five to 10 years,” former Bureau of Labor Statistics chief Erika McEntarfer told the Times. “People are looking to data to answer that question, and it’s just too difficult.”
Coders, learn to philosophize. (New York Times) Trouble landing a tech job? Consider becoming a (very specific kind) of philosopher. “I think the demand for philosophers with A.I. training is, if anything, outstripping the supply right now,” says NYU philosophy professor David Chalmers. “It’s an area I encourage students to go into.” Indeed, this small cohort has seen their job prospects skyrocket as frontier labs have rushed to employ people capable of engaging in the thorniest questions around AI’s impact on humanity and the possibility of AI consciousness. (For more on AI and philosophy, check out our AI lab philosopher draft from April.)
We build AI tools for readers like you. Write brilliantly with Spiral. Organize files automatically with Sparkle. Deliver yourself from email with Cora. Dictate effortlessly with Monologue. Collaborate with agents on documents with Proof.
For sponsorship opportunities, reach out to [email protected].
Sign in to read for free. We'll also send you a prompt that assesses your skills and gives you practical next steps.
Continue with GoogleView all login optionsThe entire economy is up for grabs in a post-AI world. What will you build?
Elbow Grease is a new accelerator from Gutter Capital in NYC. This is not a finishing school for fundraising—it’s where you build a business with people who’ve done it before. The program offers a $300,000 initial investment, weekly coaching from Gutter partners Dan Teran and James Gettinger, and 1:1 mentorship from a Series B+ or exited founder. Come work among the Gutter portfolio and learn from industry veterans Scott Belsky, Gokul Rajaram, and Hunter Walk, to name a few. Apply by July 31!