Today ainotis Join

My notis

ResearchPublished ExplainedAll news from that day

How a mixture-of-experts model activates only part of its parameters

Mistral says Large 4 has 1 trillion parameters but 49 billion active ones; the Switch Transformers and Mixtral papers describe how a mixture of experts saves compute, not memory.

Share

Check our sources · 11 facts from 4 sources
Diagram titled Mixture of Experts Layer: inputs enter a router, which sends them to two of several stacked expert blocks; gating weights combine their outputs.Albert Q. Jiang et al., Mistral AI / Mixtral of Experts, Figure 1: Mixture of Experts Layer / CC BY 4.0, cropped
Figure 1 of the Mixtral of Experts paper: a router picks two of eight experts for each input. source · CC BY 4.0

Key points

  1. In Mixtral, a router picks two of 8 experts for each token at every layer and combines their outputs.
  2. Mixtral has 46.7 billion parameters in total but uses 12.9 billion per token, so Mistral says it runs at the speed and cost of a 12.9 billion model.
  3. The saving is in compute, not memory: the Mixtral paper says memory costs for serving are proportional to its 47 billion total parameters.

What happened

In an ordinary model, every parameter takes part in every token. The Switch Transformers paper puts it this way: models typically reuse the same parameters for all inputs, and a mixture of experts instead selects different parameters for each incoming example.

The result, the authors write, is a sparsely-activated model with a very large number of parameters but a constant computational cost. Mistral's Mixtral model shows how it is built.

Each layer has 8 feedforward blocks, called experts, and for every token at every layer a router network picks two of them and combines their outputs. The router is a gating network: Mixtral takes a softmax over the top-K logits of a linear layer, so experts whose gate is zero need not be computed at all.

That is why the paper separates two counts.

The total or sparse parameter count grows with the number of experts, while the active parameter count, the parameters used for one token, grows with K. Mixtral has 46.7 billion parameters in total and uses 12.9 billion per token, and Mistral says it therefore runs at the speed and cost of a 12.9 billion model.

The saving is in compute, not in memory: the paper says the memory cost of serving Mixtral is proportional to its 47 billion total parameters. The Mixtral authors also looked for experts that specialise in a subject and did not find obvious patterns by topic.

Mistral Large 4 applies the same idea at a larger scale: Mistral describes it as a 1 trillion-parameter model with 49 billion active parameters, and as a mixture of experts. Mistral has not yet published the architecture; it says it will give details when it releases the weights.

What it means for you

Our view

When a vendor lists a total and an active parameter count, the active count is what sets the compute for each token, and the total is what sets the memory needed to serve the model, according to the Mixtral paper.

Mistral describes Large 4 as 1 trillion parameters with 49 billion active, so if it follows Mixtral, its speed and cost could be compared with a much smaller model while serving it still needs room for the full count. Mistral has not yet published the architecture and says it will share details with the weights.

Before you compare mixture-of-experts models on price, ask the vendor which count the price and the hardware requirement follow.

It adds no new facts.

Share this story

Your reaction

We count reactions per story and day, never who reacted. The counts help us choose what goes in the monthly issue. If you are signed in, your own page shows yours too.

Check our sources

Every sentence above is checked against these 4 sources.

1 Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient SparsityarXiv · 11 Jan 2021 · 2 facts Open the source
  1. The Switch Transformers paper says models typically reuse the same parameters for all inputs, and that Mixture of Experts instead selects different parameters for each incoming example. Quote: "models typically reuse the same parameters for all inputs. Mixture of Experts (MoE) defies this and instead selects different parameters for each incoming example."

    In deep learning, models typically reuse the same parameters for all inputs. Mixture of Experts (MoE) defies this and instead selects different parameters for each incoming example.
  2. The Switch Transformers paper says the result is a sparsely-activated model with outrageous numbers of parameters but a constant computational cost. Quote: "The result is a sparsely-activated model -- with outrageous numbers of parameters -- but a constant computational cost."

    The result is a sparsely-activated model -- with outrageous numbers of parameters -- but a constant computational cost.
2 Mixtral of ExpertsarXiv · 8 Jan 2024 · 5 facts Open the source
  1. The Mixtral paper says each layer of Mixtral 8x7B is composed of 8 feedforward blocks (experts) and that for every token, at each layer, a router network selects two experts to process the current state and combine their outputs. Quote: "each layer is composed of 8 feedforward blocks (i.e. experts). For every token, at each layer, a router network selects two experts to process the current state and combine their outputs."

    each layer is composed of 8 feedforward blocks (i.e. experts). For every token, at each layer, a router network selects two experts to process the current state and combine their outputs.
  2. The Mixtral paper says its gating is the softmax over the Top-K logits of a linear layer, and that if the gating vector is sparse the outputs of experts whose gates are zero can be avoided. Quote: "If the gating vector is sparse, we can avoid computing the outputs of experts whose gates are zero."

    If the gating vector is sparse, we can avoid computing the outputs of experts whose gates are zero. There are multiple alternative ways of implementing G(x) [6, 15, 35], but a simple and performant one is implemented by taking the softmax over the Top-K logits of a linear layer [28].
  3. The Mixtral paper distinguishes the total parameter count, which grows with the number of experts n, from the active parameter count used to process an individual token, which grows with K up to n. Quote: "the model’s total parameter count (commonly referenced as the sparse parameter count), which grows with n n , and the number of parameters used for processing an individual token (called the active parameter count), which grows with K K up to n n ."

    This motivates a distinction between the model’s total parameter count (commonly referenced as the sparse parameter count), which grows with n, and the number of parameters used for processing an individual token (called the active parameter count), which grows with K up to n.
  4. The Mixtral paper says the memory costs for serving Mixtral are proportional to its sparse parameter count, 47B. Quote: "The memory costs for serving Mixtral are proportional to its sparse parameter count, 47B"

    The memory costs for serving Mixtral are proportional to its sparse parameter count, 47B, which is still smaller than Llama 2 70B.
  5. The Mixtral paper reports that it did not observe obvious patterns in the assignment of experts based on topic. Quote: "Surprisingly, we do not observe obvious patterns in the assignment of experts based on the topic."

    Surprisingly, we do not observe obvious patterns in the assignment of experts based on the topic.
3 Mixtral of expertsMistral AI · 1 fact Open the source
  1. Mistral says Mixtral has 46.7B total parameters but only uses 12.9B parameters per token, and so processes input and generates output at the same speed and for the same cost as a 12.9B model. Quote: "Mixtral has 46.7B total parameters but only uses 12.9B parameters per token. It, therefore, processes input and generates output at the same speed and for the same cost as a 12.9B model."

4 Mistral Large 4Mistral AI · 6 Oct 2026 · 3 facts Open the source
  1. Mistral describes Mistral Large 4 as a 1 trillion-parameter natively multimodal model with 49 billion active parameters. Quote: "ML4 is a 1 trillion-parameter natively multimodal model with 49 billion active parameters."

    ML4 is a 1 trillion-parameter natively multimodal model with 49 billion active parameters.
  2. Mistral lists Mistral Large 4 as an open-weight hybrid instruct-and-reasoning MoE. Quote: "Open-weight hybrid instruct-and-reasoning MoE with multimodal input"

    Open-weight hybrid instruct-and-reasoning MoE with multimodal input
  3. Mistral says that, as it works toward releasing the weights, it will share further details on the model architecture. Quote: "As we work toward releasing the weights, we will share further details on the model architecture"

    As we work toward releasing the weights, we will share further details on the model architecture, additional benchmarks, and our post-training methodology.

Topics

The morning email

On the mornings we publish: the three top stories and up to four short ones. Free.

We email you a link to confirm. An issue may include one sponsor, always labelled Sponsored · Advertisement. Our emails count opens and clicks, not who made them. Unsubscribe in one click. What we keep