Home / Compare models
Compare AI models by what they are actually good at
No benchmark tables, because benchmark tables go stale and were never measuring your work anyway. What is here is what each model is reached for, where it costs you, and how to put two of them on the same question and read the disagreement yourself.
Claude vs GPT
The most asked comparison in AI, and the one where a single verdict helps least. They are close enough that the difference shows up in your task rather than in a benchmark table.
Claude vs Gemini
A genuine split rather than a ranking. One is built to be careful and one is built to be quick, and which is better depends entirely on whether you are going to act on the answer immediately.
Gemini vs GPT
For most people this comparison is settled by something other than quality: one has a free tier with no card, the other is the model everything else already integrates with.
DeepSeek vs GPT
The interesting question is not which is better on a single answer. It is which lets you run the model ten times, or on every commit, or over the whole document, without thinking about the bill.
Claude vs DeepSeek
Both are reached for when thinking matters more than speed. They differ in what they cost to run and in how much of the thinking they let you see.
Llama vs Mistral
The comparison that matters if you have decided nothing is leaving your computer. Both run locally through Ollama, both cost nothing per token, and the trade is size against efficiency.
The models
One line each on what people reach for them for.
| Model | Made by | Reached for | Free route |
|---|---|---|---|
| Claude | Anthropic | Long-form writing, code review, and work where following a complicated instruction exactly matters more than answering fast | Paid, through an Anthropic key |
| GPT | OpenAI | Breadth: the widest tooling, the most integrations, and the model most third-party software assumes you are using | Paid through OpenAI, or free through GitHub Models |
| Gemini | Very long inputs and multimodal work, and for anyone who wants a capable model without a card on file | Free tier, no card, at aistudio.google.com | |
| DeepSeek | DeepSeek | Reasoning and code at a fraction of what the frontier labs charge, which changes what is worth running at all | Paid, but cheap enough that most people never notice the bill |
| Llama | Meta | Running a capable model on hardware you control, with no per-token bill and nothing leaving the machine | Free locally through Ollama, or free hosted through Groq |
| Mistral | Mistral AI | Efficiency: strong output per unit of compute, and a European vendor for teams that need one | Free locally through Ollama, or a Mistral key |
Or stop reading comparisons and run one
Agent Mesh puts several models on the same question, lets them read and criticise each other, and converges on one answer with the disagreement left visible. On your own keys, five of which are free to obtain.
Get a free key Read published runs