Checked fact 7197 Oct 2026Research
Mistral says Mixtral has 46.7B total parameters but only uses 12.9B parameters per token, and so processes input and generates output at the same speed and for the same cost as a 12.9B model. Quote: "Mixtral has 46.7B total parameters but only uses 12.9B parameters per token. It, therefore, processes input and generates output at the same speed and for the same cost as a 12.9B model."
The exact words it rests on
Concretely, Mixtral has 46.7B total parameters but only uses 12.9B parameters per token. It, therefore, processes input and generates output at the same speed and for the same cost as a 12.9B model.
What the source said when we opened it, on 7 Oct 2026.
The source
Mixtral of experts
Checked
Checked by the notis newsroom on , against the source above.
In the story
How a mixture-of-experts model activates only part of its parameters 7 Oct 2026
Cite this fact
Anyone may quote this address. It does not change; if we correct the story, this page says so.