| Field | Value |
|---|---|
| Type | AI research company |
| Country | China |
| Parent | High-Flyer (quantitative hedge fund) |
| Key strength | Compute efficiency; open-weight releases |
DeepSeek is a Chinese AI company, backed by the quantitative hedge fund High-Flyer, that builds frontier-competitive models with substantially less compute than US counterparts, using architectural innovations (MoE, MLA, GRPO) and efficiency-focused engineering. It is frequently cited as a case study in whether compute-poor labs can compete with frontier US companies, in fast-follow dynamics, in distillation, and in US-China AI competition.
Snapshot
Valuation
| Date | Valuation | Detail | Source |
|---|---|---|---|
| 2026-07-16 | ~$52B (implied, filing) | Chinese corporate filing implies a valuation around $52B — consistent with the first round's post-money, predating any second-round close | reuters.com |
| 2026-07-14 | ~$71B pre-money (talks) | Preliminary talks for a second funding round, 37% above the ~$52B post-money of the first round; capital earmarked for gigawatt-scale data centers and in-house inference chips | techtimes.com |
| 2026-06-03 | Just under $60B | Round being finalized at over 50 billion yuan (US$7.4B); reported as six times the April mark | scmp.com |
| 2026-06-03 | $52–59B | Maiden external round slated to raise roughly $7 billion at this valuation | reuters.com |
Revenue
| Date | Figure | Detail | Source |
|---|---|---|---|
| 2026-07-15 | ~$500M annualized | Annualized revenue nearing $500 million; company preparing financial statements by December 2026 for a mainland-China IPO targeted for 2027 | theinformation.com |
Funding rounds
| Date | Amount | Round / Event | Source |
|---|---|---|---|
| 2026-07-25 | Second round suspended | Sought ≥10B yuan (~$1.5B) at $71B; verbally told prospective investors it would not sign agreements in the coming days, after viral circulation of an unverified transcript attributed to Liang Wenfeng | fortune.com |
| 2026-06-03 | ~$7.4B (50B+ yuan) | Round being finalized; committed: Tencent (10B yuan), CATL (5B), NetEase and JD.com (3B each); founder Liang Wenfeng putting in ~20 billion yuan of his own capital, the bulk self-funded | scmp.com |
| 2026-06-03 | ~$7B raise | Maiden external round slated to raise roughly $7 billion; prices the round first reported May 8 | reuters.com |
| 2026-05-08 | up to ~$7.35B (50B yuan) | First-ever outside funding round, reported as the largest by a Chinese AI company on record; founder-CEO Liang Wenfeng writing the biggest check himself | theinformation.com |
Overview
DeepSeek was founded in 2023 and is headquartered in China. It is a subsidiary of the quantitative hedge fund High-Flyer, which long supplied its capital and compute. Its models are designed to be frontier-competitive while using roughly 10× less training compute than Anthropic or OpenAI (Source: epochai.substack.com), with open-weight releases a recurring feature of its strategy. The company is led by founder-CEO Liang Wenfeng.
Models
| Model | Release | Key Feature | |
|---|---|---|---|
| [[deepseek-v4\ | DeepSeek V4 Pro / V4 Flash]] | Apr 24, 2026 | 1.6T MoE / 49B active (Pro), 284B / 13B (Flash); 1M context; CSA architecture; cheapest in tier |
| [[deepseek-v3\ | DeepSeek-V3]] | Dec 2024 | MoE: 671B params / 37B active; efficiency exemplar |
| [[deepseek-r1\ | DeepSeek-R1]] | Jan 2025 | Open-weights reasoning; briefly matched o1; triggered stock market reaction |
DeepSeek-R1, released January 2025, was an open-weights reasoning model that briefly matched OpenAI's o1 and prompted a stock market reaction. DeepSeek-V3 (December 2024) was a 671B-parameter mixture-of-experts model activating 37B parameters per token, treated as an exemplar of compute efficiency.
DeepSeek released V4 Pro (1.6T MoE / 49B active) and V4 Flash (284B / 13B) on April 24, 2026, both at 1M-token context, with new Compressed Sparse Attention and Heavily Compressed Attention architectures. V4 Pro was priced at $1.74/$3.48 per million tokens and V4 Flash at $0.14/$0.28, described as the cheapest in their respective classes. SemiAnalysis assessed V4 Pro as 3–6 months behind the US frontier (Source: simonwillison.net; newsletter.semianalysis.com). Full coverage is at DeepSeek V4 Pro / V4 Flash.
Writing in MIT Technology Review on April 24, 2026, Caiwei Chen framed the V4 release along three axes: a roughly 60% lower cost than V3 driven by 1.6T-MoE compression and DSA attention; validation of Chinese chip-supplier viability, with Huawei Ascend training compatibility disclosed for the first time as a deliberate design constraint; and acceleration of the Chinese open-weight ecosystem against Reflection, Llama, Qwen, and Mistral. Chen characterized the release as a policy-relevant event rather than merely a model release (Source: technologyreview.com).
The pricing posture reversed later in the year. Bloomberg reported on August 13, 2026 that DeepSeek is raising prices for its flagship V4 models by more than four times, effective August 16, 2026, ahead of a possible public offering, and that its rates remain below those of its largest rivals even after the increase (Source: bloomberg.com). The increase follows the V4-Pro-0813 release the previous day and roughly reverses the permanent cut described below, which had taken V4-Pro to a quarter of its prior rate.
The change also replaces flat API pricing with time-of-day rates. From 16:00 UTC on August 16, 2026, V4-Pro moves from $0.435 per million cache-miss input tokens and $0.87 per million output tokens to $0.66 and $1.98 off-peak, and $1.32 and $3.96 at peak. Reuters reported the changes represent increases ranging from 50% to more than 1,100% depending on model, token category and time of use (Source: venturebeat.com). Peak-versus-off-peak differentiation is a departure from the flat per-token pricing that characterized the V4 line at launch; see Inference Economics and Token Pricing.
On May 23, 2026, DeepSeek said it would make permanent a 75% price cut on its flagship V4-Pro model, cutting V4-Pro API costs to 0.025–6 yuan per million tokens (~$0.0035–$0.83), down from 0.1–24 yuan. The company declined to say whether the move reflected improved supply of Huawei Ascend 950 chips, leaving open whether the cut was driven by a reduction in compute cost or by competitive price pressure in the Chinese model market. The cut continued the V4-line pricing posture, V4 Pro and V4 Flash having already been the cheapest in their respective tiers at the April launch (Source: reuters.com). See Inference Economics and Token Pricing.
Key innovations
- MLA (Multi-head Latent Attention): a more efficient attention variant (Source: epochai.substack.com)
- GRPO: a reinforcement-learning approach that Epoch AI described as demonstrating "great research taste" (Source: epochai.substack.com)
- MoE at scale: activating only 37B of 671B parameters per token, a constraint-driven efficiency approach
Technical papers
- DeepSeek-V3 Technical Report — arXiv (Dec 2024): 671B MoE, FP8 training, DualPipe parallelism.
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning — Nature (Jan 2025): reasoning emerges from pure reinforcement learning.
Funding and valuation
DeepSeek had previously eschewed external funding, relying on parent High-Flyer. In May 2026 it began its first-ever outside funding round, reported on May 8 as a raise of up to roughly $7.35B (50B yuan) and described as the largest by a Chinese AI company on record, with founder-CEO Liang Wenfeng writing the biggest check himself. The Information reported that the round had prompted DeepSeek to expedite commercialization and accelerate its model-release cadence toward industry norms (Source: theinformation.com). Reuters described the shift to outside money, given the round's size, as a signal that the cost of a frontier-competitive Chinese model had grown beyond what hedge-fund cash flow could sustain even with DeepSeek's compute-efficiency advantages.
By June 3, 2026, Reuters reported the maiden external round was slated to raise roughly $7 billion at a valuation of $52–59 billion (Source: reuters.com). The same day, the South China Morning Post reported the round being finalized at over 50 billion yuan (US$7.4B) and a valuation just under $60B, six times the April mark, with committed investors Tencent (10B yuan), CATL (5B), and NetEase and JD.com (3B each), and Liang Wenfeng putting in roughly 20 billion yuan of his own capital, leaving the bulk of the round self-funded (Source: scmp.com). CATL, the Chinese EV battery maker, was first reported on May 22, 2026 to be planning to invest in the round, widening DeepSeek's backer base beyond Liang Wenfeng and High-Flyer into China's industrial-manufacturing capital (Source: theinformation.com). The round closed on June 16, 2026, raising more than 50 billion yuan (about $7.4 billion) at a valuation exceeding $50 billion, under an unusual structure in which only a Chinese state entity received voting rights (Source: reuters.com).
On July 14, 2026, DeepSeek entered preliminary talks for a second funding round at an approximately $71 billion pre-money valuation, 37 percent above the roughly $52 billion post-money figure from the first round, with the new capital earmarked for gigawatt-scale data centers and in-house inference chips (Source: techtimes.com). The Information reported the company's annualized revenue was nearing $500 million as of July 15, 2026 (Source: theinformation.com). The Wall Street Journal reported on July 14 that DeepSeek is preparing to list shares in Shanghai in 2027, seeking funds for research in its competition with Anthropic and other U.S. developers; the company is preparing financial statements by December 2026 for the offering (Source: wsj.com; theinformation.com; bloomberg.com). Founder Liang Wenfeng, who retains roughly 78 percent of the company, saw his net worth reach about $36 billion on July 14, 2026 (Source: techtimes.com).
The second round was suspended on July 25, 2026. DeepSeek verbally told prospective investors it would not sign agreements in the coming days; it had sought at least 10 billion yuan (about $1.5 billion) at a $71 billion valuation. The suspension followed the viral circulation of a transcript attributed to Liang discussing reliance on Nvidia chips and China's lag behind the United States — comments whose authenticity has not been verified. The same report described the first round as having closed in June 2026 "at $7 billion," a figure that does not reconcile with the roughly $52 billion post-money valuation implied by the July 16 Chinese corporate filing and most likely refers to the amount raised rather than the valuation; the report does not specify which (Source: fortune.com).
Competitive position
DeepSeek trains its models with roughly 10× less compute than Anthropic or OpenAI (Source: epochai.substack.com). The Stanford HAI AI Index 2026 placed the US-China capability gap at 2.7% on benchmarks as of March 2026, with models trading the lead multiple times (Source: Stanford HAI AI Index Report 2026). Zhipu AI CEO Tang Jie offered a contrary read, stating: "The truth may be that the gap is actually widening" (Source: epochai.substack.com). Founder Liang Wenfeng gave his own assessment in a May 20, 2026 investor call whose leaked transcript circulated on July 23, 2026: in the 3-hour-44-minute session he said China's gap with the U.S. is primarily compute resources rather than talent, defended open-sourcing as DeepSeek's core strategy, and predicted consolidation among AI labs (Source: luizasnewsletter.com; aiproem.substack.com). He also gave concrete chip-access figures: DeepSeek sought 200,000 Huawei 950 chips for a frontier training run but could obtain only 16,000, against Huawei's expected 750,000-chip output for 2026 (Source: fredgao.com; transformernews.ai).
DeepSeek recurs across several competition and governance debates:
- Fast-follow. DeepSeek-R1 caught up to OpenAI's o1 within months with less training compute, cited as a case study of the fast-follow problem.
- Export controls. DeepSeek's progress despite chip restrictions is cited against the argument that export controls create durable advantages, though Epoch AI argues efficiency cannot fully bridge a 10× compute gap (Source: epochai.substack.com).
- Distillation. Anthropic accused DeepSeek of distilling from Claude's outputs; MiniMax may have obtained roughly 100B tokens from Claude interactions (Distillation).
- Supply-chain designation. In the Anthropic-DoW conflict, DeepSeek had not been designated a supply-chain risk despite links to the Chinese military, while Anthropic, an American company, was threatened with that designation.
Reuters reported on July 7, 2026, citing three people familiar with the matter, that DeepSeek is developing its own AI chip to reduce reliance on both Nvidia and Huawei silicon (Source: reuters.com). The move parallels Zhipu AI's reported inquiries about a bespoke processor the same week. See Semiconductor Supply Chain, Export Controls (AI).
Government and national-security response
The V4 release on April 24, 2026 coincided with US federal actions framing DeepSeek as a national-security concern. The Office of Science and Technology Policy issued an NSTM-4 memo titled "Adversarial Distillation of American AI Models" (see Adversarial Distillation). The State Department sent a global cable ordering posts to spotlight alleged DeepSeek IP theft (Source: reuters.com). On April 22, 2026, Tencent and Alibaba together committed over $20B to DeepSeek-centric infrastructure builds (Source: bloomberg.com).
See also
Related pages: (Source: epochai.substack.com), China and the US Are Running Different AI Races, The Least Understood Driver of AI Progress, Clawed, Distillation, DeepSeek-V3 Technical Report, DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning, DeepSeek V4 Pro / V4 Flash, Adversarial Distillation.