DeepSeek R1 Shakes the AI Model Market
Published:
DeepSeek released R1 on January 20, 2025, alongside a detailed technical report describing the model’s architecture and training methodology. DeepSeek is a Chinese AI research lab founded in 2023 as an offshoot of High-Flyer Capital Management, a quantitative hedge fund, with the stated goal of pursuing AGI research independent of product revenue pressures. The R1 model itself was a 671-billion-parameter Mixture-of-Experts architecture with approximately 37 billion parameters active per token during inference — a design that substantially reduced the computational cost of a forward pass relative to a dense model of equivalent total parameter count. The training pipeline centered on GRPO (Group Relative Policy Optimization), a reinforcement learning technique that rewarded correct answers on verifiable tasks (mathematics, competitive programming, formal logic) without requiring a separately trained reward model for each capability area. An intermediate model, R1-Zero, was trained using only RL with no supervised fine-tuning whatsoever; it spontaneously developed extended internal reasoning traces — long chains of intermediate steps visible in the model’s output before its final answer — a behavior the researchers described as emerging from the optimization objective rather than being explicitly trained. The final R1 model incorporated a small amount of supervised fine-tuning on high-quality examples followed by additional RL training, and performed comparably to OpenAI’s o1 model on benchmarks including AIME 2024 (mathematics olympiad problems), Codeforces (competitive programming), and MATH-500.
The release terms were what made R1 globally significant beyond its benchmark scores. DeepSeek released the model weights under the MIT License, permitting commercial use, modification, and redistribution. They also released distilled versions — smaller models derived by training Qwen-2.5 and Llama 3 series models to imitate R1’s extended reasoning behavior — in sizes from 1.5B to 70B parameters. These distilled models, particularly DeepSeek-R1-Distill-Qwen-32B and R1-Distill-Llama-70B, offered strong reasoning performance on consumer hardware. DeepSeek’s API pricing for R1 was approximately $0.55 per million input tokens and $2.19 per million output tokens — roughly 20-30× cheaper than OpenAI’s o1 at comparable benchmark performance. DeepSeek had previously released DeepSeek-V3 (the base non-reasoning model) on December 26, 2024, reporting a training cost of approximately $5.6 million — a figure that attracted widespread attention given estimates that frontier models from US labs cost $50–100M or more to train, though the comparison involved different hardware generations and assumptions.
The market reaction was immediate and sharp. On January 27, 2025, NVIDIA’s stock fell approximately 17 percent in a single trading session — wiping roughly $600 billion in market capitalization — as investors reassessed assumptions about GPU demand for AI training. The concern was that if competitive reasoning models could be trained at a fraction of the cost attributed to US frontier models, the thesis driving the AI hardware buildout might be less robust than believed. Within weeks, DeepSeek’s iOS app briefly ranked as the most-downloaded app in the United States App Store, surpassing ChatGPT. The release intensified several ongoing technical debates: whether the reported training cost was comparable to the full cost (including prior research compute for architecture decisions), whether export controls on NVIDIA’s H100/H800 chips had meaningfully constrained or inadvertently accelerated efficiency research at Chinese labs, and whether open-weight models with strong reasoning capability changed the calculus for enterprise AI deployment. For developers, the distilled R1 variants provided the first practical path to running o1-grade reasoning behavior on locally hosted hardware without API dependency, accelerating adoption of quantized inference runtimes like llama.cpp, vLLM, and Ollama for production workloads.
