
Meta’s Llama family is the best-known series of open-weight language models: you can download the model files and run them on your own computer or server. Here is what each generation brought, in plain English.
In this article
Why Llama matters
- Open weights: companies can run Llama privately, fine-tune it, and avoid sending data to a third party.
- An ecosystem: tools like Ollama and llama.cpp, and thousands of fine-tuned variants, grew around it.
- Free to try: this site’s AI Helper runs Llama models through Cloudflare’s free tier.
The timeline
| Release | Sizes | What was new |
|---|---|---|
| Llama 1 — Feb 2023 | 7B–65B | Research-only licence. Showed a 13B model could rival much larger ones when trained on more data. Its weights leaked and kick-started the open-model community. |
| Llama 2 — Jul 2023 | 7B, 13B, 70B | Licence allowed most commercial use. Chat-tuned versions released. 4K context. |
| Llama 3 — Apr 2024 | 8B, 70B | Much better quality from far more training data (15T+ tokens) and a larger tokenizer. 8K context. |
| Llama 3.1 — Jul 2024 | 8B, 70B, 405B | 128K context, multilingual, and the 405B “frontier-class” open model; better tool calling. |
| Llama 3.2 — Sep 2024 | 1B, 3B; 11B/90B vision | Tiny models for phones and laptops, plus the first Llama models that can read images. |
| Llama 3.3 — Dec 2024 | 70B | Quality close to 3.1 405B in a 70B model — cheaper and faster to run. |
| Llama 4 — Apr 2025 | Scout, Maverick (Behemoth previewed) | Mixture-of-experts design (only part of the model is active per token), native text+image input, and very long context on Scout. |
Mixture of experts in one paragraph
Llama 4 Scout is described as 17B active parameters with 16 experts. The model contains many “expert” sub-networks; for each token a router picks a few of them. You get the knowledge of a large model while only computing a fraction of it per token — faster and cheaper, but it still needs enough memory to hold every expert.
Which one should you use?
| Need | Pick |
|---|---|
| Laptop, offline, private | Llama 3.2 3B or 3.1 8B with Ollama |
| Best free quality for text | Llama 3.3 70B (cloud) — the AI Helper default |
| Images plus text, long documents | Llama 4 Scout |
| Fine-tuning for a narrow task | A small 3.x model with LoRA |
Limitations to keep in mind
- Knowledge cut-off: each model only knows what was in its training data; it does not know newer events.
- Hallucination: like every LLM it can state wrong facts confidently — see why.
- Licence: the Llama licence is permissive but not a standard open-source licence; very large companies and some uses have conditions. Read it before shipping a product.
- Hardware: bigger models need serious GPUs; quantized versions trade a little quality for much less memory.
Compare Llama with other families in the big AI model comparison.
✨ Ask AI about this article
Stuck on a step? Ask a question and the AI answers using this article.
Free · AI can be wrong