AI Models Compared: GPT, Claude, Gemini, Llama, DeepSeek, Qwen, Mistral and More

⏱ 3 min readUpdated 28 September 2026

There are now dozens of capable AI models. This guide groups the popular families, what each is known for, and how to try them for free. Model versions change quickly — use it as a map, then check each provider’s current documentation.

In this article
  1. Closed (hosted) vs open-weight
  2. The main families
  3. How to read benchmarks
  4. Common limitations (all models)
  5. Which to choose?

Closed (hosted) vs open-weight

  • Closed models run only on the provider’s servers (ChatGPT, Claude, Gemini). Usually the most capable; you pay per use or by subscription.
  • Open-weight models can be downloaded and run anywhere (Llama, gpt-oss, Qwen, DeepSeek, Mistral, Gemma, GLM, Kimi). Good for privacy, cost control and fine-tuning.

The main families

Family Maker Known for Free way to try
GPT / o-series / gpt-oss OpenAI All-round quality, reasoning models, huge ecosystem ChatGPT free tier; gpt-oss in our AI Helper
Claude Anthropic Careful writing, long documents, strong coding and agentic work claude.ai free tier
Gemini / Gemma Google Very long context, multimodal (images, video, audio); Gemma is the open-weight sibling Gemini app; Gemma in the AI Helper
Llama Meta The most popular open-weight family — history AI Helper, Ollama
DeepSeek DeepSeek (China) Strong reasoning (R1) at low training cost; open weights DeepSeek chat; R1 distill in the AI Helper
Qwen Alibaba Wide size range, strong coding (Qwen Coder), multilingual Qwen Chat; Qwen Coder in the AI Helper
Mistral Mistral AI (France) Efficient models, European provider Le Chat; Mistral Small in the AI Helper
Kimi Moonshot AI Large mixture-of-experts models, long context, coding/agent tasks Kimi chat app
GLM Zhipu / Z.ai Open-weight models aimed at reasoning and agents Z.ai chat

How to read benchmarks

Every launch comes with benchmark charts. Know what they measure:

Benchmark Measures
MMLU / MMLU-Pro Multiple-choice knowledge across ~57 school and professional subjects
GPQA Very hard, “Google-proof” graduate-level science questions
HumanEval, LiveCodeBench Writing code that passes tests
SWE-bench Verified Fixing real GitHub issues in real projects — closest to real developer work
AIME, MATH Competition mathematics
LMArena (Chatbot Arena) People vote blind between two answers — measures what users prefer
⚠️ Benchmarks can be “trained to” and don’t cover your task. The only benchmark that matters is 20 of your real questions — try them on two or three models and compare.

Common limitations (all models)

  • Knowledge cut-off dates; no awareness of recent events unless connected to search.
  • Hallucinations, especially with numbers, names and citations.
  • Context limits: very long documents can be truncated or partly ignored.
  • Usage caps on free tiers; privacy terms differ — never paste confidential data into a free public chatbot without checking.

Which to choose?

Task Good starting point
Excel formulas, quick answers Llama 3.3 70B or GPT-OSS (free in our AI Helper)
VBA / SQL / Python Qwen Coder, Claude, GPT
Long reports, careful writing Claude, Gemini
Private, offline Llama or Qwen with Ollama
Hard maths/logic Reasoning models (o-series, DeepSeek R1, GPT-OSS 120B)
✨ Ask AI about this article

Stuck on a step? Ask a question and the AI answers using this article.

Free · AI can be wrong

Leave a Reply

Your email address will not be published. Required fields are marked *