Skip to content
The Shape of Intelligence

DeepSeek-V3 trained for $5.6 million

A Chinese hedge fund's laboratory releases a 671-billion-parameter open model that matches GPT-4o, trained on export-restricted chips for a reported fraction of the usual cost.

category
model
significance
3 of 5
people
Liang Wenfeng
organisations
DeepSeek

what had to happen · 38 events back to 1943

Every event this one built on, transitively, in order. Direct influences are marked.

DeepSeek was the AI laboratory of High-Flyer, a quantitative hedge fund in Hangzhou, and on 26 December 2024 it released the weights of a model that few outside China had been watching for. DeepSeek-V3 was a mixture-of-experts transformer with 671 billion parameters, of which 37 billion were active for any token, trained on 14.8 trillion tokens. It matched GPT-4o and Claude 3.5 Sonnet on most benchmarks. The technical report put the final training run at 2.8 million GPU-hours on NVIDIA H800s, the chip designed to comply with American export controls, and about $5.6 million at rental prices.

The figure excluded research, failed runs and the cluster itself, and it was still an order of magnitude below the estimates for comparable Western models. The report explained how: an attention variant that compressed the memory of the context, a load-balancing scheme for the experts, training in eight-bit floating point, and hand-tuned communication that got around the restricted chips' slower interconnect.

The release was a month before DeepSeek-R1 turned the same base into a reasoning model and moved the markets. V3 is on this timeline as the evidence that the frontier could be approached with far less money, and from outside the three laboratories, than anyone had assumed.

what it led to · 4 events downstream, through 2025

Built on it directly:

  1. 2025DeepSeek-R1VII

And, through them, by era:

sources · 1

See this era in the exhibition →Back to the timeline