GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series with 320B total parameters and 18B active parameters. It incorporates several architectural improvements over GLM-5.2 including a hybrid architecture, sharply reducing long-context serving costs while preserving precise long-context and Manifold-Constrained Hyper-Connections (mHC) to further improve scaling efficiency.
Model availability and access depend on your workspace and plan.
About the model
GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series with 320B total parameters and 18B active parameters. It incorporates several architectural improvements over GLM-5.2 including a hybrid architecture, sharply reducing long-context serving costs while preserving precise long-context and Manifold-Constrained Hyper-Connections (mHC) to further improve scaling efficiency.
What you get back. The response appears directly in your Relam chat, where you can continue the conversation.
Capabilities & limits
Ask questions, give instructions and continue with follow-up messages in a conversation.
The catalogue lists image understanding as a supported model capability. Available attachment controls depend on your workspace.
The model can work with tools as part of a response. Relam determines which tools are available for your request and enforces their permissions.
GLM 5.3 Flash has a 1,048,576-token context window. This is the model’s capacity for the text it considers in a request, including prompts and conversation context. Tokens are pieces of text, rather than a word or page count. The amount of context sent by your workspace can be lower.
The catalogue sets a maximum output of 4,096 tokens for GLM 5.3 Flash. This is the response budget, distinct from its context window. Responses can be shorter depending on your request and the limits applied by your workspace.
Questions & answers
GLM 5.3 Flash is listed as a reasoning-capable model. Its configured reasoning levels are low, high, max. The default level is low. Available controls depend on your workspace.