DeepSeek, a Chinese artificial intelligence (AI) model developer, is drawing attention after its newly unveiled lightweight model surpassed the performance of its own flagship model through post-training alone. Analysts say the achievement breaks the conventional belief that more parameters necessarily mean better model performance.

According to Chinese IT media outlet TMTPost on the 2nd, DeepSeek released the application programming interface (API) of the official version of its new AI model "DeepSeek V4 Flash" on the 31st of last month. This comes about four months after the company released preview versions in April of V4's lightweight model "Flash" and its flagship model "Pro," respectively, and marks the launch of the final edition of Flash. DeepSeek said it also plans to release the official version of Pro as soon as possible.
This model added only post-training while maintaining the same architecture and parameter count as the existing Flash preview. In theory, the smaller the model scale, the lower the performance should be. Yet it overwhelmed the "V4 Pro Preview," which has roughly eight times as many parameters, shocking the industry. The buzz grew further when Tesla founder Elon Musk, who has praised the pace of Chinese AI development, began following DeepSeek's official account on X (formerly Twitter) the day after the announcement.
In fact, this model scored 82.7 on the "Terminal-Bench 2.1" benchmark, which evaluates coding ability, far exceeding the V4 Pro Preview's 72.4. On the Artificial Analysis Intelligence Index, which aggregates nine metrics including Terminal-Bench 2.1, it recorded 50 points, ahead of the Pro Preview by 6 points. This is a level similar to Google's recently unveiled "Gemini 3.6 Flash." Analysts say this is the result of DeepSeek intensively boosting the capabilities needed for complex agent tasks, such as code execution and autonomous decision-making, during the post-training stage. TMTPost assessed that "once model performance exceeds a certain level, post-training may be more important than scale." Its price is also markedly low at $0.28 per one million output tokens, compared not only with Claude Opus 4.8 ($25) but also with "Luna" ($1.2), the low-cost model of OpenAI's GPT-5.6.
Chinese companies have recently unveiled a string of high-performance, low-cost models, putting the U.S. industry on edge. A prime example is "Kimi K3," a next-generation open-weight model unveiled by Moonshot AI last month. Kimi K3 recorded 57 points on the Artificial Analysis Intelligence Index, ranking third with a gap of just 2 to 3 points from top U.S. models "Claude Fable 5" (60 points) and "GPT-5.6 Sol" (59 points). Some assessed that, with Kimi K3, the technology gap with U.S. models has narrowed to about 2 to 3 months. GLM-5.2, unveiled by Zhipu AI early last month, also drew attention by implementing performance close to that of top U.S. models at about one-fifth of the price.






