Checked fact 7177 Oct 2026Research
The Mixtral paper says its gating is the softmax over the Top-K logits of a linear layer, and that if the gating vector is sparse the outputs of experts whose gates are zero can be avoided. Quote: "If the gating vector is sparse, we can avoid computing the outputs of experts whose gates are zero."
The exact words it rests on
If the gating vector is sparse, we can avoid computing the outputs of experts whose gates are zero. There are multiple alternative ways of implementing G(x) [6, 15, 35], but a simple and performant one is implemented by taking the softmax over the Top-K logits of a linear layer [28].
What the source said when we opened it, on 7 Oct 2026.
The source
Mixtral of Experts
Checked
Checked by the notis newsroom on , against the source above.
In the story
How a mixture-of-experts model activates only part of its parameters 7 Oct 2026
Cite this fact
Anyone may quote this address. It does not change; if we correct the story, this page says so.