NVIDIA DGX Station GB300 compared with 4x / 8x RTX PRO 6000 Server edition GPU Server

The NVIDIA DGX Station GB300 is an interesting offering for teams requiring cloud independence. It is primarily intended for office deployment, as a relatively quiet machine with a 1600W power supply unit.

However, some teams are also able to host their own servers in a datacenter. Therefore, it is useful to compare the NVIDIA DGX Station GB300 with a powerful server with several NVIDIA RTX PRO 6000 Blackwell Server edition GPUs.

Executive Summary: DGX Station GB300 vs. RTX Pro 6000 Server Edition GPUs

For serving frontier models to a workgroup, the DGX Station GB300 is the right choice.

The DGX Station GB300 outperforms 4x RTX PRO 6000 Blackwell Server Edition in the relevant performance measures (FP4, FP8, FP16/BF16) and features higher memory bandwidth for a slightly higher price.

A significant advantage is the bigger block of single GPU memory, reducing setup work and communication latencies for sharding models.

Additionally, the DGX Station has a significantly lower power draw compared to the servers with 4x and 8x RTX Pro, making it compatible with standard office power sockets. The lower power draw means lower noise, and a lower electricity bill.

A server with 8x RTX PRO 6000 Server edition will be able to serve models at higher precision. Additionally, it features an x86 architecture compatible with legacy software, and more CPU RAM. If used correctly, the higher aggregate GPU memory bandwidth could support higher token/s throughput on it. Comparing it with a dual NVIDIA DGX Station GB300 Setup is advisable.

We can support you in finding the right choice for your application, reach out to us:

Detailed Comparison

The DGX Station, and the multi-GPU servers are available from several vendors. Supermicro currently leads with competitive pricing, therefore we will use a Supermicro example. Supermicro, next to HP, also offers an optional rack mounting solution for the DGX Station.

NVIDIA DGX Station GB300Server with 4x RTX PRO 6000 Server EditionServer with 8x RTX PRO 6000 Server Edition
Base model / chassisSupermicro Super AI Station ARS-511GD-NB-LCCSupermicro AS-5126GS-TNRT bareboneSupermicro AS-5126GS-TNRT barebone
Form factordesktop tower
optional rack mounting kit (5U)
5U rack-mountable server5U rack-mountable server
CPU1x NVIDIA Grace 72-core Neoverse V2 CPU (ARM-based)2 x AMD Turin 9555 DP/UP 64C/128T 3.2G 256MB 360W SP5 HF
(EPYC 9005-series, Supermicro SKU: PSE-TUR9555-1142)
2 x AMD Turin 9555 DP/UP 64C/128T 3.2G 256MB 360W SP5 HF
(EPYC 9005-series, Supermicro SKU: PSE-TUR9555-1142)
CPU RAM496GB ECC LPDDR5X768GB
= 24 * 32GB DDR5-6400 2Rx8 (16Gb) ECC RDIMM
768GB
= 24 * 32GB DDR5-6400 2Rx8 (16Gb) ECC RDIMM
CPU memory bandwidth396 GB/s1,23 TB/s = 2 * 614GB/s per CPU1,23 TB/s = 2 * 614GB/s per CPU
CPU & GPU memory unified
(coherent memory)
yes, via NVLink ®-C2C, at 900 GB/s interconnect bandwidthnono
GPU(s)1x NVIDIA Blackwell Ultra GPU4x NVIDIA RTX PRO 6000 Blackwell Server Edition, PCIe 58x NVIDIA RTX PRO 6000 Blackwell Server Edition, PCIe 5
GPU RAM252GB ECC HBM3eper GPU: 96GB GDDR7
aggregate: 384GB GDDR7
per GPU: 96GB GDDR7
aggregate: 768GB GDDR7
GPU memory bandwidth7,1 TB/sper GPU: 1.597 GB/s
aggregate: 6,4 TB/s
per GPU: 1.597 GB/s
aggregate: 12,8 TB/s
multi-instance GPU (MIG)yes, up to 7 MIGsyes, up to 4 MIGs per GPUyes, up to 4 MIGs per GPU
High-speed Ethernet portNVIDIA ConnectX-8 SuperNIC, dual port QSFP112ConnectX-6 Dx 100GbE Ethernet Adapter Card, dual port QSFP56ConnectX-6 Dx 100GbE Ethernet Adapter Card, dual port QSFP56
Hi-speed Ethernet bandwidth 800 Gbit/s
= 2x 400 Gbit/s
200 Gbit/s
= 2x 100 Gbit/s
200 Gbit/s
= 2x 100 Gbit/s
Storage4x M.2 Slots:
2x 2TB M.2 PCIe 5.0×4 NVMe boot drives
1x 2TB M.2 data drive
2x SSD 2.5″ SATA 480GB 1DWPD TLC E, SED/TCG 7mm2x SSD 2.5″ SATA 480GB 1DWPD TLC E, SED/TCG 7mm
Power consumptionup to 1600W
About 4000W total
Up to 360W per CPU
Up to 600W per GPU
About 7000W total
Up to 360W per CPU
Up to 600W per GPU
office suitabilitywell suited:
acceptable noise level – liquid cooling system
can be powered through office power socket
desktop tower form factor
not suitable:
noisy
requires industrial strength power sockets
rack form factor, requires rack
not suitable:
noisy
requires industrial strength power sockets
rack form factor, requires rack
price (net, USD)ca. $105.000ca. $100.000ca. $150.000

AI performance: 1x DGX Station GB300 vs. 4 / 8x NVIDIA RTX PRO 6000 Blackwell Server edition

NVIDIA DGX Station GB300RTX PRO 6000 Server Edition, per GPU4x RTX PRO 6000 Server Edition8x RTX PRO 6000 Server Edition
CUDA parallel processing coresTBD24.06496.256192.512
NVIDIA Tensor CoresTBD752
(fifth-generation)
3.0086.016
NVIDIA RT CoresN/A188
(fourth-generation)
7521.504
FP4 PFLOPS20 PFLOPS
15 PFLOPS without sparsity
4 PFLOPS
1871,2 TFLOPS without sparsity
16 PFLOPS
7,5 PFLOPS without sparsity
32 PFLOPS
15 PFLOPS without sparsity
FP810 PFLOPS
5.000 TFLOPS without sparsity
2 PFLOPS
935.6 TFLOPS without sparsity
8 PFLOPS
3.742 TFLOPS without sparsity
16 PFLOPS
7.485 TFLOPS without sparsity
FP6 Tensor Core10 PFLOPSN/AN/AN/A
Int8 Tensor Core330 TOPS (? NVIDIA)
5.000 TOPS without sparsity (? Google)
935,6 TOPS without sparsity 3.742 TOPS without sparsity3.742 TOPS without sparsity
FP16/BF16 Tensor Core5 PFLOPS
2.500 TFLOPS without sparsity
1 PFLOP
467,8 TFLOPS without sparsity
4 PFLOPS
1.871 TFLOPS without sparsity
8 PFLOPS
3.742 TFLOPS without sparsity
TF32 Tensor Core2.5 PFLOPS
1.250 TFLOPS without sparsity
234 TFLOPS without sparsity936 TFLOPS without sparsity1.872 TFLOPS without sparsity
FP3280 TFLOPS120 TFLOPS480 TFLOPS960 TFLOPS
FP64 / FP64 Tensor Core1,3 TFLOPS1,8 TFLOPS7,2 TFLOPS14,4 TFLOPS
RT Core performanceN/A355 TFLOPS1.420 TFLOPS2.840 TFLOPS
GPU memory252 GB HBM3e memory

can use CPU’s 496 GB LPPDR5X via NVLink-C2C

= 748 GB RAM total
96 GB GDDR7 with ECC384 GB 768 GB
Memory interfaceHBM3e / LPDDR5X512-bit GDDR7GDDR7GDDR7
Memory bandwidth7,1 TB/s for 252GB HBM3e
396 GB/s for 496GB LPDDR5X
1.597 GB/s6,4 TB/s aggregate12,8 TB/s aggregate
Power consumptionsystem: up to 1.600WGPU: up to 600W GPUs: up to 2.400WGPUs: up to 4.800W

For LLM inference, there are two distinct phases:

  • The prefill phase (loading the user’s prompt is compute bound
  • The decode phase (generating tokens) is memory bound

As we can see, the DGX Station GB300 has more compute power than 4 x RTX PRO 6000 Server Edition GPUs for FP4, FP8, FP16/BF16. FP4 and FP8 performance are especially important for LLM inference.

However, even a single RTX PRO 6000 outperforms the DGX Station GB300 Superchip for FP32, and FP64 performance. Additionally, it has Raytracing capabilitie, which the DGX Station GB300 lacks.

Strictly from the compute power perspective, 8x RTX PRO 6000 Blackwell Server Edition outperform the DGX Station GB300 for FP4, FP8, FP16/BF16, and of course also for FP32 and FP64. This, however, requires a task which allows you to load the GPUs evenly with processing power.

For scientific workloads, the compute precision you need matters – for high precision scientific, or rendering workloads, the RTX PRO 6000 Server Edition cards are a better match.

For the decode phase, memory bandwidth matters – here the DGX Station GB300 also outperforms 4x RTX PRO 6000 GPUs, and is the clear winner – expected to deliver better token/s performance.

For 8x RTX PRO 6000 GPUs, the situation is more complex. The aggregate memory bandwidth is higher, hower, the GPUs might stall due to required inter-GPU communication over the PCI Express link. With the proper engineering, throughput could be increased, compared to the DGX Station GB300.

2x DGX Station GB300 are a strong contender for a comparison. 2x DGX Station GB300 will have an even higher aggregate memory bandwidth of 14,2 TB/s, higher compute performance, lower overall power draw, and less model shards.

pi3g can support you to understand the right fit for vLLM inference servers.

Note:

The Int8 performance value for the DGX Station GB300 given in NVIDIA’s official datasheet does not jibe with the FP8, FP16 progression. I would expect the performance to be similar to FP8. 330 TOPS is at the performance level of an edge accelerator card. The value from Google, 5.000 TFLOPS (sic) seems more reasonable, but should probably read 5.000 TOPS, since we are dealing with Int8 here.

References:

Model fit, scalability and performance

The DGX Station GB300 has a bigger and faster single block of memory – 252GB ECC HBM3e at 7,1 TB/s memory bandwidth, compared with a single NVIDIA RTX PRO 6000 Blackwell Server edition GPU.

Even when combining four RTX PRO 6000 GPUs, their aggregate memory bandwidth is still lower compared to the DGX Station GB300’s GPU memory bandwidth.

Depending on your application and model, you are able to optimize for different things.

The following models are known to work on the DGX Station GB300:

model namesizearchitectureprecisionDGX Station1x RTX Pro 60004x RTX Pro 60008x RTX Pro 6000
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF1630B/3BMoENVFP4yesyes yesyes
Step 3.7 – Flash198B/11BMoEBF16yes (FP8)noyes (FP8, EP)yes (EP)
DeepSeek-V4-Flash284B/13BMoEFP8yes (FP4)noyes (FP8, EP)yes (FP8, EP)
Qwen3.6-27B27BDenseBF16yesyesyesyes
Qwen3.6-35B-A3B35B/3BMoEBF16yes yesyesyes
MiniMax-M2.7230B/10BMoEFP8yes (FP8)noyes (FP8, EP)yes (FP8, EP)
gemma-4-26B-A4B-it26 B/4BMoEBF16yesyesyesyes
NVIDIA-Nemotron-3-Super-120B-A12B-BF16120B/12BMoEBF16yesyes (FP4)yes (EP)yes (EP)

Note: the context can be put into CPU memory for the DGX Station, due to the unified memory design. EP in the table stands for Expert parallelism.

Expert Parallelism (EP)

Today’s frontier open weight models predominantly use a mixture of experts architecture (MoE). In this architecture, a router sends tokens to a small number of selected experts.

In the case of the DGX Station, all models in the list above fit completely into the fast GPU memory, using the specified precision. The router does not have any additional PCIe latency, the tokens do not need to travel to another GPU for computation.

Additionally, no special sharding setup is needed for the models, making management and development easier.

For RTX Pro 6000 GPUs, experts will need to be distributed, with typically several experts being distributed to each GPU. Special attention needs to be paid to the experts fitting, and additional latencies introduced by the inter-GPU communication during inference.

Having up to 768 GB aggregate GPU memory on 8 x RTX PRO 6000 GPUs will allow you to load models with higher precision. The DGX Station has 748 GB memory in total, however it’s CPU memory is 4x slower than a single RTX PRO 6000’s memory. The bandwidth matters more.

Data parallelism

For data parallelism, the same model runs on several GPUs, data is processed in parallel on these GPUs.

As 4x RTX PRO 6000 underperform the DGX Station GB300 on every relevant measure (compute for FP4, FP8, FP16/BF16, and memory bandwidth), and there is additional PCIe overhead, you would shoot yourself in the foot trying to obtain data parallelism this way.

The aggregate performance and memory bandwidth of 8x RTX PRO 6000 are higher. In the right circumstances, this can lead to higher token/s throughput. However, it should be carefully compared to the performance of 2x DGX Station GB300 Systems, especially from the perspective of reduced power usage.

Tensor parallelism

In tensor paralellism, every GPU holds a fraction of every layer, and computes a partial result, that has to be recombined. The necessary synchronisation between the GPUs is very data-intensive.

Typically, NVLink and direct GPU-to-GPU communication is required for this to perform well.

Since the RTX PRO 6000 does not have an NVLink option, we do not recommend to attempt tensor paralellism on them.

Pipeline Parallelism

Different layers live on different GPUs, which are chained together like an assembly line.

This is a good fit for multi-GPU clusters, like 8x RTX PRO 6000 Blackwell, and could lead to this GPU cluster outperforming the DGX Station GB300.

Redundancy

The DGX Station GB300 has a single power supply, a single GPU, and a single CPU. These are single points of failure. A server system with 4x or 8x RTX Pro 6000 Blackwell GPUs will be more robust against failure, and degrade gracefully instead of failing completely.

x86 / CPU / CPU RAM

A final consideration for our comparison is the CPU architecture of the systems.

The DGX Station uses ARM cores. With x86 being the de-facto standard in the server world, some legacy applications would need to be recompiled, or ported. Compatibility might be an issue.

The AMD EPYC used in this server has a higher memory bandwidth – per CPU, and, of course, together for both.

Depending on the tasks, additional processing could be loaded into the CPU cores, and be the better overall solution. Therefore, we recommend a holistic look at your application, to determine the right fit.