Today ainotis Join

My notis

ModelsPublished Top storyAll news from that day

AWS lists Z.ai's GLM 5.3 as generally available on Bedrock

AWS says Z.ai's GLM 5.3 is generally available on Amazon Bedrock to eligible enterprise customers, in a What's New entry dated 5 October 2026.

Share

Check our sources · 5 facts from 1 source
AWS blog launch card reading Introducing GLM 5.3 on Amazon Bedrock, with a green-to-purple gradient strip down its right edge.Image: Amazon Web Services
source · The company's own image of its product, used to report on it

Key points

  1. AWS describes GLM 5.3 as a mixture-of-experts model with 753B total parameters and roughly 40B active per token.
  2. AWS says it has a 1-million-token context window and up to 128K output tokens, with reasoning always enabled and selectable effort levels.
  3. AWS says it is available to eligible enterprise customers through the US and Global cross-Region inference profiles, and Bedrock support includes explicit prompt caching.

What happened

AWS says Bedrock now supports GLM 5.3, a mixture-of-experts model with 753 billion total and about 40 billion active parameters per token, a 1-million-token context window and up to 128K output tokens. Bedrock support includes explicit prompt caching. It is available to eligible enterprise customers through the US and Global cross-Region inference profiles.

What it means for you

Our view

These are AWS's own descriptions of the model's size and limits, and they say nothing yet about how it performs on your tasks. The selectable effort levels mean each request trades latency and token use against task performance, so cost per task will depend on the level you choose.

Explicit prompt caching on system prompts and messages matters most where the same context is sent repeatedly. Availability is limited to eligible enterprise customers, so access in your own account is the first question.

If you run workloads on Bedrock, check whether your account is eligible and which inference profile applies, then test cost and latency at each effort level on your own prompts.

This is our view of the facts above. It adds no new facts.

Share this story

Your reaction

Each tap adds one to the count. We count reactions per story and day, never who reacted. The counts help us choose what goes in the monthly issue. If you are signed in to My notis, your own page shows your reactions too.

Check our sources

We checked every sentence above against this source (5 facts in all).

1 GLM 5.3 by Z.ai is now generally available on Amazon BedrockAmazon Web Services · 5 Oct 2026 · 5 facts Open the source
  1. AWS's What's New entry, dated 5 October 2026, is titled "GLM 5.3 by Z.ai is now generally available on Amazon Bedrock".

    GLM 5.3 by Z.ai is now generally available on Amazon Bedrock Posted on: Oct 5, 2026
  2. AWS describes GLM 5.3 as a mixture-of-experts model with 753B total parameters and roughly 40B active per token.

    GLM 5.3 is Z.ai’s flagship model, a mixture-of-experts architecture with 753B total parameters and roughly 40B active per token
  3. AWS says GLM 5.3 has a 1-million-token context window and up to 128K output tokens, with reasoning always enabled and selectable effort levels.

    It combines a 1-million-token context window with up to 128K output tokens, and reasoning is always enabled with selectable effort levels so you can trade latency and token consumption against task performance.
  4. AWS says Bedrock support includes explicit prompt caching with cache points on system prompts and messages.

    Bedrock support for GLM 5.3 includes explicit prompt caching with cache points on system prompts and messages, helping you reduce latency and input costs when reusing context across model calls.
  5. AWS says GLM 5.3 "is available to eligible enterprise customers" and is accessible through the US and Global cross-Region inference profiles.

    GLM 5.3 is available to eligible enterprise customers, and is accessible through the US and Global cross-Region inference profiles.

We link every source we used.

Topics

The morning email

On the mornings we publish, usually soon after 07:00 Oslo time: the day's three top stories, what they mean for you, and up to four short news items. Free.

We email you a link to confirm. An issue may include one sponsor, always labelled Sponsored · Advertisement. Our emails count opens and clicks, not who made them. Unsubscribe in one click. What we keep