Companies ainotis Join
My notis
Company

Mistral

2 stories · first 7 Oct · latest 7 Oct

Products and models

Facts we checked

Each fact was checked against the page it links to.

  • Mistral says that, as it works toward releasing the weights, it will share further details on the model architecture. Quote: "As we work toward releasing the weights, we will share further details on the model architecture"

    Mistral Large 4 · Mistral AI · archived copy · checked 7 Oct 2026 · Cite this fact · the story

  • Mistral lists Mistral Large 4 as an open-weight hybrid instruct-and-reasoning MoE. Quote: "Open-weight hybrid instruct-and-reasoning MoE with multimodal input"

    Mistral Large 4 · Mistral AI · archived copy · checked 7 Oct 2026 · Cite this fact · the story

  • Mistral describes Mistral Large 4 as a 1 trillion-parameter natively multimodal model with 49 billion active parameters. Quote: "ML4 is a 1 trillion-parameter natively multimodal model with 49 billion active parameters."

    Mistral Large 4 · Mistral AI · archived copy · checked 7 Oct 2026 · Cite this fact · the story

  • The Mixtral paper reports that it did not observe obvious patterns in the assignment of experts based on topic. Quote: "Surprisingly, we do not observe obvious patterns in the assignment of experts based on the topic."

    Mixtral of Experts · arXiv · archived copy · checked 7 Oct 2026 · Cite this fact · the story

  • The Mixtral paper says the memory costs for serving Mixtral are proportional to its sparse parameter count, 47B. Quote: "The memory costs for serving Mixtral are proportional to its sparse parameter count, 47B"

    Mixtral of Experts · arXiv · archived copy · checked 7 Oct 2026 · Cite this fact · the story

  • Mistral says Mixtral has 46.7B total parameters but only uses 12.9B parameters per token, and so processes input and generates output at the same speed and for the same cost as a 12.9B model. Quote: "Mixtral has 46.7B total parameters but only uses 12.9B parameters per token. It, therefore, processes input and generates output at the same speed and for the same cost as a 12.9B model."

    Mixtral of experts · Mistral AI · archived copy · checked 7 Oct 2026 · Cite this fact · the story

  • The Mixtral paper distinguishes the total parameter count, which grows with the number of experts n, from the active parameter count used to process an individual token, which grows with K up to n. Quote: "the model’s total parameter count (commonly referenced as the sparse parameter count), which grows with n n , and the number of parameters used for processing an individual token (called the active parameter count), which grows with K K up to n n ."

    Mixtral of Experts · arXiv · archived copy · checked 7 Oct 2026 · Cite this fact · the story

  • The Mixtral paper says its gating is the softmax over the Top-K logits of a linear layer, and that if the gating vector is sparse the outputs of experts whose gates are zero can be avoided. Quote: "If the gating vector is sparse, we can avoid computing the outputs of experts whose gates are zero."

    Mixtral of Experts · arXiv · archived copy · checked 7 Oct 2026 · Cite this fact · the story

  • The Mixtral paper says each layer of Mixtral 8x7B is composed of 8 feedforward blocks (experts) and that for every token, at each layer, a router network selects two experts to process the current state and combine their outputs. Quote: "each layer is composed of 8 feedforward blocks (i.e. experts). For every token, at each layer, a router network selects two experts to process the current state and combine their outputs."

    Mixtral of Experts · arXiv · archived copy · checked 7 Oct 2026 · Cite this fact · the story

  • Mistral describes Large 4 as an open-weight hybrid instruct-and-reasoning MoE with multimodal input. Quote: "Open-weight hybrid instruct-and-reasoning MoE with multimodal input"

    Mistral Large 4 · Mistral AI · archived copy · checked 7 Oct 2026 · Cite this fact · the story

  • Mistral says a significant share of Large 4's training data was multilingual, spanning more than 160 languages, including every official language of the European Union. Quote: "a significant share of ML4’s training data was multilingual, spanning more than 160 languages, including every official language of the European Union."

    Mistral Large 4 · Mistral AI · archived copy · checked 7 Oct 2026 · Cite this fact · the story

  • Mistral says that on Lakera's public B3 AI Security Benchmark Large 4 resists 93.3% of attacks. Quote: "On Lakera’s public B3 AI Security Benchmark , ML4 resists 93.3% of attacks"

    Mistral Large 4 · Mistral AI · archived copy · checked 7 Oct 2026 · Cite this fact · the story

  • Mistral reports Large 4 scores 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA and 28.3% on Terminal-Bench 4. Quote: "scoring 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, and 28.3% on Terminal-Bench 4."

    Mistral Large 4 · Mistral AI · archived copy · checked 7 Oct 2026 · Cite this fact · the story

  • Mistral says that on the Artificial Analysis Cyber Index Large 4 ranks among the top five models globally; this is Mistral's own account of the result. Quote: "it ranks among the top five models globally and leads open-weight models developed outside China by a wide margin."

    Mistral Large 4 · Mistral AI · archived copy · checked 7 Oct 2026 · Cite this fact · the story

  • Mistral says Large 4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in its own datacenters in Europe, and the public preview is served on the same infrastructure. Quote: "ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own datacenters in Europe. The public preview is served on that same infrastructure."

    Mistral Large 4 · Mistral AI · archived copy · checked 7 Oct 2026 · Cite this fact · the story

  • Until the weights are released, Mistral says it is red-teaming the model with cybersecurity leaders, vetted partners and state authorities, who access the same model with reduced moderation and expanded cyber capabilities. Quote: "we are red-teaming the model in real-world settings with cybersecurity leaders, vetted partners, and state authorities, who will access the same model with reduced moderation and expanded cyber capabilities."

    Mistral Large 4 · Mistral AI · archived copy · checked 7 Oct 2026 · Cite this fact · the story

  • Mistral says it will release the weights by the end of the month. Quote: "We will release the weights by the end of the month."

    Mistral Large 4 · Mistral AI · archived copy · checked 7 Oct 2026 · Cite this fact · the story

  • Mistral lists the preview API price as $1.36 per million input tokens and $4.18 per million output tokens. Quote: "Input (/M tokens) $1.36 Output (/M tokens) $4.18"

    Mistral Large 4 · Mistral AI · archived copy · checked 7 Oct 2026 · Cite this fact · the story

  • Mistral describes Mistral Large 4 as a 1 trillion-parameter natively multimodal model with 49 billion active parameters. Quote: "ML4 is a 1 trillion-parameter natively multimodal model with 49 billion active parameters."

    Mistral Large 4 · Mistral AI · archived copy · checked 7 Oct 2026 · Cite this fact · the story

  • Mistral AI published Mistral Large 4 on 6 October 2026. Quote: "Introducing Mistral Large 4"

    Mistral Large 4 · Mistral AI · archived copy · checked 7 Oct 2026 · Cite this fact · the story

All 2 stories

Cases

Ongoing case: Frontier model API prices

Built from our published stories and the facts we checked, and kept up to date by them. How we correct.