Model Details

Domain:

Language

Task:

Language modeling

Code generation

Model Access:

Open weights (non-commercial)

Citations:

16922

Introduction

We introduce LLaMA, a collection of foundation language models ranging from 7B to 65B parameters. We train our models on trillions of tokens, and show that it is possible to train state-of-the-art models using publicly available datasets exclusively, without resorting to proprietary and inaccessible datasets. In particular, LLaMA-13B outperforms GPT-3 (175B) on most benchmarks, and LLaMA-65B is competitive with the best models, Chinchilla-70B and PaLM-540B. We release all our models to the research community.

Benchmarking

FLOPs

7.8e+22

Notes: 1T tokens * 13B parameters * 6 FLOP/token/parameter = 7.8e22 from paper, Llama-13B took 135,168 GPU hours using A100s 312 trillion * 135,168 * 3600 = 1.518e23 FLOPs at full utilization This implies that the actual utilization was: MFU = 7.8e22/1.518e23 = 0.514

Training

Training Code Accessibility

non-commercial license: https://docs.google.com/forms/d/e/1FAIpQLSfqNECQnMkycAp2jP4Z9TFX0cGR4uf7b_fBxjY_OjhJILlKGA/viewform

Hardware

NVIDIA A100

Size Notes: Table 2

Parameters

13000000000

Notes: 13.0B

Authors

Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, Guillaume Lample

Related Models

LLaMA-13B - Use Model

LLaMA-13B - Use Model

Model Details

Introduction

Benchmarking

Training

Parameters

Authors

Related Models

Llama 4 Behemoth preview

Llama 4 Maverick

Llama 4 Scout

Llama 3.3 70B

LLaMA-13B - Use Model

LLaMA-13B - Use Model

Model Details

Introduction

Benchmarking

Training

Parameters

Authors

Related Models

Llama 4 Behemoth preview

Llama 4 Maverick

Llama 4 Scout

Llama 3.3 70B