
There are now dozens of capable AI models. This guide groups the popular families, what each is known for, and how to try them for free. Model versions change quickly — use it as a map, then check each provider’s current documentation.
In this article
Closed (hosted) vs open-weight
- Closed models run only on the provider’s servers (ChatGPT, Claude, Gemini). Usually the most capable; you pay per use or by subscription.
- Open-weight models can be downloaded and run anywhere (Llama, gpt-oss, Qwen, DeepSeek, Mistral, Gemma, GLM, Kimi). Good for privacy, cost control and fine-tuning.
The main families
| Family | Maker | Known for | Free way to try |
|---|---|---|---|
| GPT / o-series / gpt-oss | OpenAI | All-round quality, reasoning models, huge ecosystem | ChatGPT free tier; gpt-oss in our AI Helper |
| Claude | Anthropic | Careful writing, long documents, strong coding and agentic work | claude.ai free tier |
| Gemini / Gemma | Very long context, multimodal (images, video, audio); Gemma is the open-weight sibling | Gemini app; Gemma in the AI Helper | |
| Llama | Meta | The most popular open-weight family — history | AI Helper, Ollama |
| DeepSeek | DeepSeek (China) | Strong reasoning (R1) at low training cost; open weights | DeepSeek chat; R1 distill in the AI Helper |
| Qwen | Alibaba | Wide size range, strong coding (Qwen Coder), multilingual | Qwen Chat; Qwen Coder in the AI Helper |
| Mistral | Mistral AI (France) | Efficient models, European provider | Le Chat; Mistral Small in the AI Helper |
| Kimi | Moonshot AI | Large mixture-of-experts models, long context, coding/agent tasks | Kimi chat app |
| GLM | Zhipu / Z.ai | Open-weight models aimed at reasoning and agents | Z.ai chat |
How to read benchmarks
Every launch comes with benchmark charts. Know what they measure:
| Benchmark | Measures |
|---|---|
| MMLU / MMLU-Pro | Multiple-choice knowledge across ~57 school and professional subjects |
| GPQA | Very hard, “Google-proof” graduate-level science questions |
| HumanEval, LiveCodeBench | Writing code that passes tests |
| SWE-bench Verified | Fixing real GitHub issues in real projects — closest to real developer work |
| AIME, MATH | Competition mathematics |
| LMArena (Chatbot Arena) | People vote blind between two answers — measures what users prefer |
⚠️ Benchmarks can be “trained to” and don’t cover your task. The only benchmark that matters is 20 of your real questions — try them on two or three models and compare.
Common limitations (all models)
- Knowledge cut-off dates; no awareness of recent events unless connected to search.
- Hallucinations, especially with numbers, names and citations.
- Context limits: very long documents can be truncated or partly ignored.
- Usage caps on free tiers; privacy terms differ — never paste confidential data into a free public chatbot without checking.
Which to choose?
| Task | Good starting point |
|---|---|
| Excel formulas, quick answers | Llama 3.3 70B or GPT-OSS (free in our AI Helper) |
| VBA / SQL / Python | Qwen Coder, Claude, GPT |
| Long reports, careful writing | Claude, Gemini |
| Private, offline | Llama or Qwen with Ollama |
| Hard maths/logic | Reasoning models (o-series, DeepSeek R1, GPT-OSS 120B) |
✨ Ask AI about this article
Stuck on a step? Ask a question and the AI answers using this article.
Free · AI can be wrong