DeepSeek's V4 Pro AI Model Underwhelms on Benchmarks, Excels in Cybersecurity
DeepSeek, a Chinese AI startup, has discreetly launched its updated flagship model, DeepSeek-V4-Pro-0813, which has garnered a mix of responses. Researchers in specialized fields such as cybersecurity find the model noteworthy, yet initial benchmark assessments reveal its inability to compete with leading counterparts, disappointing some developers regarding its performance and pricing. On the Artificial Analysis Intelligence Index, DeepSeek-V4-Pro-0813 achieved a score of 53, equaling Zhipu AI's GLM-5.2 from June, but falling short by four points compared to OpenAI's GPT-5.6 Terra and seven points behind Moonshot AI's Kimi K3. The Vals Index from Vals AI in San Francisco placed it 12th, lagging behind OpenAI's GPT-5.5 and advanced systems like Kimi K3 and Anthropic's Claude Opus 5. Vals AI highlighted on Wednesday that the model had particular difficulties with tasks in a sandboxed terminal and creating intricate financial models in Excel. Nonetheless, its cybersecurity performance remains a noted advantage as DeepSeek navigates the competitive AI arena against established entities like OpenAI and Anthropic.
Key facts
- DeepSeek released DeepSeek-V4-Pro-0813, an updated flagship model.
- The model scored 53 on the Artificial Analysis Intelligence Index.
- It tied with Zhipu AI's GLM-5.2 but trailed OpenAI's GPT-5.6 Terra and Moonshot AI's Kimi K3.
- On the Vals Index, it ranked 12th, behind GPT-5.5 and frontier models like Kimi K3 and Claude Opus 5.
- Vals AI reported struggles with sandboxed terminal tasks and Excel financial models.
- The model impressed researchers in cybersecurity.
- Some developers were disappointed with its overall capabilities and pricing.
- The release is part of DeepSeek's competition with other AI startups.
Entities
Institutions
- DeepSeek
- Zhipu AI
- OpenAI
- Moonshot AI
- Vals AI
- Anthropic
- Artificial Analysis
Locations
- China
- San Francisco
- United States
- Beijing