We use analytics and advertising tools by default. You can update this anytime.
A practical guide to commodity intelligence, owning your AI stack, and deciding when to still reach for the frontier
Not long ago, the best frontier models were as unreliable as they were impressive. They could explain quantum mechanics and write decent code, but couldn’t resist fabricating quotes or miscounting the Rs in “strawberry.” Anyone doing serious work used the strongest model they could access.
Since then, the best models have raced ahead in specialized areas like research math, advanced cybersecurity, and scientific discovery—far beyond what most of us need for everyday work. Meanwhile, they’ve grown more expensive to use and more locked down.
But the same progress has also turned AI models into commodities. Capabilities that once required a frontier model are now available in cheaper, widely available models. For routine tasks, today’s open models are competitive on both price and quality, and companies have begun to use them alongside frontier models from OpenAI and Anthropic.
It doesn’t take a genius to summarize meeting notes—though it does take one to discover new antibiotics. Most work falls somewhere in between, and the practical question to ask is, “Which model is good enough for my task?”
I’ve been tinkering with open models since Meta released Llama in 2023. Back then, open models were research curiosities. Today, they handle a growing share of my work, from professional coding to creative projects to automating digital chores.
The cost savings are nice, but the bigger prize is independence. When an API goes down, you have a backup. When a model vanishes from the menu, you have other options. This guide walks through the open model stack one layer at a time. It shows you how to assess your needs, choose a model and host, and put it to work.
When you use ChatGPT or Claude, the app connects to a model running on the company’s servers. The company also stores your conversation history so you can pick up where you left off across devices. Even if you install the app on your computer, the model still runs remotely. The model itself is just a large file containing billions of numbers, called “weights,” that encode the patterns it learned in training.
U.S. frontier labs keep their weights secret. Open-weight models make those files available to download and run yourself, subject to their licenses. Some—but not all—open-weight models are also open source: they share training code and information about the training data, with broad rights to use, modify, and redistribute them.
Model weights are just the start. Every other layer of the product—from hosting to tools to interface—also has open alternatives.
You can decide how much of each layer to run yourself. Paying a provider saves you setup and maintenance work. Running it yourself gives you more say over where your data goes and how the system works, but you become responsible for keeping it running. It’s not all or nothing, and different tasks lend themselves to different setups.
Most of the strongest open models today come from Chinese labs. U.S. and European labs release open models too, but not as many, and the models they offer are generally not as capable.
Giving models away isn’t charity. Free has long been a competitive business strategy. Google made Android free to phone makers, helping secure a place for its search engine on mobile devices. Chinese labs publish open models for similar reasons: marketing, tapping a global research community, and pressuring U.S. rivals whose business depends on renting out closed models.
If the origin concerns you, consider that Cursor fine-tuned Kimi K2.5 into Composer 2, the model behind its coding agent, and Perplexity has long offered open models alongside frontier ones. Using a model from a Chinese lab doesn’t require using the lab’s app or sending your data to China. A model is a file of numbers. It’s auditable by independent researchers, and it can run on a host you trust.
Whether a model is good enough depends on what you ask it to do. A task gets harder when the instructions are vague, the consequences of a mistake are serious, or the model has to make many decisions without supervision. The same model can be reliable for one job and inadequate for another.
The strongest models are better at interpreting unclear requests and working through long, complicated problems. Use them when the work is hard to define in advance, takes many steps, or would be costly to get wrong.
Smaller models can handle clear, repeatable work that is easy to verify. Good instructions, relevant reference material, and checks along the way can make them reliable on harder tasks.
Start with the task. Support-ticket triage is routine and mistakes are easy to spot, making it a good candidate for an open model. A one-off payment-system refactor is complicated and costly to get wrong, so use the strongest model you can afford.
The AI you already use may know enough about your work to spot good candidates. This prompt asks it to use any memory or history it has, then fill in the gaps with a few questions.
Help me identify work I currently give to frontier AI that may be simple enough for cheaper, faster, commodity AI. Use what you already know about my projects and habits, then interview me to fill the gaps.
For background on open-weight models and the framework used here, see the full guide:
https://every.to/guides/getting-started-with-open-models
The goal is a shortlist of three small experiments, not a complete automation strategy. Prioritize work I already do with AI. You may include an adjacent task that is not currently AI-assisted when it is an unusually strong fit.
Commodity intelligence includes cheaper or open-weight models, whether remote hosted or local. Do not assume I am technical. Keep the conversation focused on what I am trying to accomplish, what goes in, what useful work comes out, and how I would know whether the result is good enough. Leave models, APIs, code, and automation platforms for a later conversation.
A strong candidate usually has several of these qualities:
Treat these as judgment criteria, not a rigid numerical score. Frequency alone does not make a task suitable. For each candidate, also test:
Do not force an entire task onto commodity intelligence when only part of it fits. Look for a safer boundary: commodity intelligence can extract, sort, summarize, format, or prepare a first pass while a person or frontier model retains the ambiguous judgment. Prefer a smaller useful experiment to an impressive autonomous system.
Keep frontier intelligence for work that is one-off, poorly defined, hard to verify, taste-dependent, consequential, or cheap only when it succeeds on the first try. When a candidate does not survive scrutiny, say so at that point in the conversation and explain why. Do not save rejected tasks for a required section in the final answer.
Once you understand enough of my work, present the result in four layers. Give me the recommendation before the audit trail.
Open with a conversational summary of the pattern you found in two to four sentences. Focus on what makes parts of my work promising for commodity intelligence, not on recapping your interview process.
Present the best three candidates in priority order as a numbered list. Give each a plain-language name followed by two or three sentences. Tell me what the experiment would accomplish, why this task appears suitable, and where human or frontier judgment should remain. Write this as advice to me, not as a form or scorecard.
Prefer a short, credible list over filling the quota with weak ideas. If fewer than three tasks genuinely fit, return fewer and explain what evidence is missing.
Follow the summary with compact supporting details for each candidate. Use the same names and order so the sections are easy to connect. Include:
Keep these details concise and avoid repeating the recommendation above. They should support it, not bury it.
Recommend one candidate to try first and explain the choice in two or three sentences. Favor easy learning and low downside over maximum theoretical savings. End by asking whether I want to refine that experiment or turn it into a lightweight evaluation plan. Do not proceed into implementation until I choose.
Upgrade your membership to continue reading the full guide.
Subscribe to continue