DeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts model with 552B backbone parameters that natively processes images and text at up to 1M token context. Its Causal Encoder-Decoder architecture activates only 8B parameters during prefill and 16B during decode, cutting the KV cache footprint to roughly a quarter of DeepSeek-V4-Flash for cost-efficient agentic workloads.
Model availability and access depend on your workspace and plan.
About the model
DeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts model with 552B backbone parameters that natively processes images and text at up to 1M token context. Its Causal Encoder-Decoder architecture activates only 8B parameters during prefill and 16B during decode, cutting the KV cache footprint to roughly a quarter of DeepSeek-V4-Flash for cost-efficient agentic workloads.
What you get back. The response appears directly in your Relam chat, where you can continue the conversation.
Capabilities & limits
Ask questions, give instructions and continue with follow-up messages in a conversation.
The catalogue lists image understanding as a supported model capability. Available attachment controls depend on your workspace.
The model can work with tools as part of a response. Relam determines which tools are available for your request and enforces their permissions.
DeepSeek V4.1 Flash has a 1,000,000-token context window. This is the model’s capacity for the text it considers in a request, including prompts and conversation context. Tokens are pieces of text, rather than a word or page count. The amount of context sent by your workspace can be lower.
The catalogue sets a maximum output of 4,096 tokens for DeepSeek V4.1 Flash. This is the response budget, distinct from its context window. Responses can be shorter depending on your request and the limits applied by your workspace.
Questions & answers
DeepSeek V4.1 Flash is listed as a reasoning-capable model. Its configured reasoning levels are none, low, high, max. The default level is low. Available controls depend on your workspace.