ainotis Join
My notis

Checked fact 7177 Oct 2026Research

The Mixtral paper says its gating is the softmax over the Top-K logits of a linear layer, and that if the gating vector is sparse the outputs of experts whose gates are zero can be avoided. Quote: "If the gating vector is sparse, we can avoid computing the outputs of experts whose gates are zero."

The exact words it rests on

If the gating vector is sparse, we can avoid computing the outputs of experts whose gates are zero. There are multiple alternative ways of implementing G(x) [6, 15, 35], but a simple and performant one is implemented by taking the softmax over the Top-K logits of a linear layer [28].

What the source said when we opened it, on 7 Oct 2026.

The source

Mixtral of Experts
arXiv · 2024-01-08

Checked

Checked by the notis newsroom on , against the source above.

In the story

How a mixture-of-experts model activates only part of its parameters 7 Oct 2026

Cite this fact

Anyone may quote this address. It does not change; if we correct the story, this page says so.