CoreWeave trains DeepSeek-V3 in two minutes in MLPerf benchmark
CoreWeave, Inc. (NASDAQ: CRWV) reported results from the MLPerf Training v6.0 benchmark suite, training the DeepSeek-V3 671B model in 2.02 minutes using 8,192 NVIDIA GB300 NVL72 GPUs across 2,048 nodes, according to a company press release.
The company submitted three GB300 NVL72 configurations on the DeepSeek-V3 671B workload, claiming the fastest results across all Closed/Available-cloud submissions in the benchmark round. At 4,096 GPUs across 1,024 nodes, training completed in 3.09 minutes, and at 2,048 GPUs across 512 nodes, the result was 5.54 minutes. CoreWeave said it was the only submitter in the v6.0 round to scale a GB300 platform beyond 2,048 GPUs on the DeepSeek-V3 workload.
On a separate configuration, CoreWeave's 4,096-GPU GB300 NVL72 deployment reached the Llama-3.1-405B reference quality target in 9.77 minutes. On an 8-node, 64-GPU NVIDIA HGX B200 cluster, the company trained GPT-OSS-20B in 26.98 minutes and Llama-3.1-8B in 16.54 minutes.
"Training DeepSeek-V3 in two minutes on the largest GB300 cluster reflects years of metal-to-model engineering investment," said Chen Goldberg, Executive Vice President of Product and Engineering at CoreWeave. "These results came from the same infrastructure our customers run in production today, not a benchmark-only setup."
CoreWeave stated the infrastructure used in the benchmark, including its networking fabric, scheduler, storage architecture, and Mission Control orchestration platform, is the same available to customers on its production cloud.
MLPerf is an industry-standard benchmarking suite used to measure machine learning training performance across hardware and software configurations.
