AWS lists Z.ai's GLM 5.3 as generally available on Bedrock
Image: Amazon Web ServicesAWS says Z.ai's GLM 5.3 is generally available on Amazon Bedrock to eligible enterprise customers, in a What's New entry dated 5 October 2026.
- AWS describes GLM 5.3 as a mixture-of-experts model with 753B total parameters and roughly 40B active per token.
- AWS says it has a 1-million-token context window and up to 128K output tokens, with reasoning always enabled and selectable effort levels.
- AWS says it is available to eligible enterprise customers through the US and Global cross-Region inference profiles, and Bedrock support includes explicit prompt caching.
AWS says Bedrock now supports GLM 5.3, a mixture-of-experts model with 753 billion total and about 40 billion active parameters per token, a 1-million-token context window and up to 128K output tokens. Bedrock support includes explicit prompt caching. It is available to eligible enterprise customers through the US and Global cross-Region inference profiles.
What it means for you Our view
These are AWS's own descriptions of the model's size and limits, and they say nothing yet about how it performs on your tasks. The selectable effort levels mean each request trades latency and token use against task performance, so cost per task will depend on the level you choose. Explicit prompt caching on system prompts and messages matters most where the same context is sent repeatedly. Availability is limited to eligible enterprise customers, so access in your own account is the first question. If you run workloads on Bedrock, check whether your account is eligible and which inference profile applies, then test cost and latency at each effort level on your own prompts.