Global Markets
Home›Global Markets›China›DeepSeek launches V4.1 Flash AI model, touting lower i…
DeepSeek launches V4.1 Flash AI model, touting lower inference costs
The new model is built on a 552 billion-parameter framework but DeepSeek says it activates 8 billion parameters for input and 16 billion for response generation.
DeepSeek has released its V4.1 Flash model, saying it outperforms its previous flagship while reducing inference costs and improving speeds, positioning the update as part of China’s push for aggressive price to performance in AI, the SCMP Economy reported.
The company said V4.1 Flash uses a new Causal Encoder Decoder architecture and is built on a mixture-of-experts design, which routes tasks to only the specific subnetworks suited to each request instead of running every query through the entire system.
DeepSeek described the model as the smallest in its new series and said it includes native multimodal visual understanding. It also said V4.1 Flash scored higher than prior models on benchmarks covering coding, cybersecurity, and autonomous agent tasks.
On Terminal-Bench 2.1, DeepSeek reported a score of 90.6 for V4.1 Flash, ahead of OpenAI’s GPT-5.6 Sol at 88.8, Moonshot AI’s Kimi K3 at 88.3, and DeepSeek’s own V4 Pro at 87.9. The company also pointed to rising hardware costs and foreign chip export curbs as reasons Chinese developers are racing to offer more efficient models.