Choosing a language model in 2026: what actually matters
New models arrive every few weeks. The leaderboard rarely decides what works in production; your tasks, your data and your constraints do.
A new language model is released every few weeks, each one topping some leaderboard. It is tempting to treat model choice as the main decision in an AI project and to switch every time the rankings move.
In my experience, the leaderboard rarely decides what works in production. Your tasks, your data and your constraints do. The model is one component of the system, and the system is what has to deliver.
Think in tiers, not names
The major providers offer families of models in roughly three tiers, and they are worth thinking about separately:
- Frontier models are the most capable, and the slowest and most expensive. Use them for complex reasoning, long multi-step agentic tasks and difficult code.
- Mid-tier models handle the bulk of production work well: drafting, summarising, extraction with judgement, most tool use.
- Small, fast models are ideal for classification, routing, simple extraction and anything you run at very high volume.
Anthropic’s Claude models, for example, come as Opus, Sonnet and Haiku; OpenAI and Google offer similar ranges. Alongside them sit open-weight families such as Llama, Mistral and Qwen, which you can host yourself.
Good systems often use several tiers at once: a small model triages and routes, and a larger one handles the cases that need real reasoning. Matching the model to the step usually saves more money than any discount negotiation.
Evaluate on your own work
Public benchmarks measure general ability. They tell you little about how a model handles your documents, your terminology and your edge cases.
What works is a small evaluation set built from real examples:
- Collect 50 to 200 real cases, including the awkward ones.
- Write down what a correct result looks like for each.
- Score each model on accuracy, the kinds of mistakes it makes, whether it follows the required format and how reliably it uses tools.
- Run the same set every time a new model comes out, or whenever you change a prompt.
An afternoon of evaluation beats a week of reading release notes.
Once that set exists, deciding whether to adopt a new model takes hours instead of weeks of debate.
The criteria that actually decide
When I compare models for a production system, these are the questions that matter:
- Quality on your evaluation set. Not in general; on your tasks.
- Tool use and instruction following. For agentic work, the model must call tools correctly, stay within its scope and stop when it should. This varies more between models than raw intelligence does.
- Latency. A model that is a little better but twice as slow can make an interactive product feel broken.
- Cost per completed task. Count the whole task, not the price per token: long prompts, retries and multiple steps add up. Prompt caching and batch processing can change the picture significantly.
- Context handling. A large context window is not the same as using that context well. Test retrieval from long inputs explicitly.
- Data protection. Where is data processed, how long is it retained, is it used for training, and is there a proper data processing agreement? For European organisations, this often narrows the choice before quality is even discussed.
- Reliability. Rate limits, regional availability and how the provider handles outages.
Hosted or open-weight?
Hosted frontier models give you the most capability with the least operational effort. Open-weight models give you control: data never leaves your infrastructure, you can fine-tune, and costs are predictable at scale.
The trade-off is operational effort and, for the hardest tasks, usually some capability. A common and sensible split is hosted models for complex reasoning and open-weight models for high-volume or particularly sensitive steps.
Design for change
The one thing I can predict is that today’s best choice will not be next year’s. So I build systems where switching models is a configuration change followed by an evaluation run, not a rewrite:
- a thin layer between the application and the model provider
- prompts and evaluation sets kept under version control, next to the code
- the model for each step set in configuration, not hard-coded
With that in place, a new release becomes an opportunity rather than a migration project.
Where I start
My default is to start each step with a strong mid-tier model, measure it on real cases, then move the hardest steps up a tier and the high-volume, simple steps down. I re-run the evaluations whenever a significant new model arrives, and only switch when the numbers say so.
The model matters. But the evaluation set, the tool design and the guardrails around it are what turn a good model into a system people can rely on.