
A server with eight NVIDIA B300 GPUs is not a “graphics card machine”: it is data center infrastructure for serving very large AI models to many users simultaneously. In this video, BIZON shows the X9000 G5, with eight NVIDIA B300 GPUs on an HGX B300, platform, running Kimi K3 with vLLM and measuring throughput for 1 to 128 concurrent users. The demo is interesting, but also a good opportunity to separate architecture, commercial server, and marketing benchmark.
Video summary: “NVIDIA B300 Server Review: 8 GPUs, 128 Users, AI Benchmarks”, from the BIZON channel, published on 7 of September 2026. The manufacturer presents the BIZON X9000 G5 with eight NVIDIA B300 GPUs on an HGX B300 platform and inference results with Kimi K3 and vLLM. Watch on YouTube.
What are NVIDIA B300, HGX B300 and BIZON X9000 G5?
The names sound similar, but they describe different layers of the solution:
- NVIDIA B300: GPU from the Blackwell Ultra generation aimed at AI, HPC, and data center inference.
- NVIDIA HGX B300: platform/baseboard with eight B300 GPUs connected via NVLink and NVSwitch.
- BIZON X9000 G5: BIZON server that integrates the HGX B300 platform in its own chassis.
- NVIDIA DGX B300: another product: a complete NVIDIA system. It is not the X9000 G5 and should not be treated as if it were the same machine.
This distinction matters when comparing CPU, system memory, network, power supplies, support, and software. Two machines can use HGX B300 but remain different products with different integration and support choices.
What the HGX B300 platform delivers
In NVIDIA's reference architecture, an HGX B300 node packs eight Blackwell Ultra GPUs. Each GPU has 288 GB of HBM3e, reaching 2,30 TB of HBM3e memory on the node. The documentation also indicates up to 8 TB/s of memory bandwidth per GPU and 64 TB/s in aggregate across the GPU set.
For large models, memory capacity is not enough: GPUs need to exchange activations, partitioned weights, and communication data without turning the external network into a bottleneck. That's why HGX B300 uses fifth-generation NVLink and NVSwitch. NVIDIA specifies 1,8 TB/s of GPU-to-GPU communication and 14,4 TB/s aggregate across the eight-GPU domain. To scale beyond a single node, the reference includes eight ConnectX-8 SuperNICs, with up to 800 Gb/s per adapter.
It's the difference between placing many independent GPUs in a server and building an accelerated domain that can cooperate when serving an LLM. To understand why memory, interconnection, and numeric format matter so much in AI, see also GPU vs TPU: how AI chips work.
What the BIZON X9000 G5 adds to the mix
The X9000 G5 is the BIZON integration in 8U form factor. In the AMD configuration listed by the manufacturer, there is support for two EPYC 9005/9004 processors, up to 3 TB of DDR5 ECC, eight front NVMe Gen5 hot-swap bays and four PCIe Gen5 x16 slots. The page also lists network options with eight OSFP ports of 800 Gb/s for InfiniBand XDR or two Ethernet links of 400 Gb/s per ConnectX-8.
There is also an Intel variant listed by BIZON, with two Xeon 6500/6700 and up to 4 TB of DDR5. Therefore, it makes no sense to talk about “the” amount of CPU RAM of the X9000 G5 without saying which configuration is being quoted. Items such as CPU, RAM, storage, adapters, and support contract are part of the project, not footnote details.
The video benchmark: Kimi K3 with vLLM
BIZON uses the vLLM as an inference server and claims to load the Kimi K3 on the eight GPU cluster. The central metric disclosed is aggregate throughput, in tokens per second, as the number of concurrent users increases:
| Concurrent users | Disclosed aggregate throughput |
|---|---|
| 1 | 99,8 tokens/s |
| 32 | 2.058,7 tokens/s |
| 64 | 3.013,4 tokens/s |
| 128 | 4.508,8 tokens/s |
These values are results reported by the manufacturer in the video for a single workload; they are not 4.508,8 tokens/s per person and do not automatically measure the experience of any application. Even so, they illustrate an important characteristic of inference servers: under concurrency, continuous batch scheduling can greatly increase the total tokens produced per second.
Why tokens per second is not enough
A throughput table is only part of the test. To decide whether a platform serves a chatbot, an agent API or an internal team, you also need to measure time to first token (TTFT), per-token latency, P95/P99 percentiles, error rate, context length, input and output size, queue policy, and actual consumption.
The video does not publish enough details for independent reproduction: exact model version and quantization, vLLM version and parameters, prompts, output limit, concurrency method, TTFT and latency percentiles are not documented in the description. Therefore, the correct reading is: a capability demonstration of the BIZON + HGX B300 + chosen workload, and not a universal ranking between GPUs or a capacity promise for all models.
How much does the BIZON X9000 G5 and each B300 cost?
Corporate AI hardware pricing is not very transparent: there is no public retail price list for B300 in the references consulted, and OEMs and integrators sell these systems through quotes. Therefore, the correct value for planning is a reference range, followed by a quote for the exact configuration.
In the query made on 11 of September of 2026, the BIZON page showed the X9000 G5 starting at US$ 496.728. It's the starting price of a complete server, not of eight standalone GPUs: CPU, system memory, NVMe, network, chassis, power supplies, integration and support also make up the proposal. Freight, taxes, import and local services may change the final cost.
As a market reference, the AI Hardware Price Index recorded in 31 August 2026 a median of US$ 544.280 per new HGX B300 node with eight GPUs, equivalent to US$ 68.035 per GPU in the index metric. In that survey, the advertisements ranged from US$ 399.999 to US$ 682.000 per node.
This figure of US$ 68.035 is neither MSRP nor the price of a standalone B300: it's the division of the advertised price of complete eight-GPU systems. The index itself notes that it works with asking prices from sellers, not completed sales, and doesn't standardize CPU, RAM, storage, or warranty. The initial price of US$ 496.728 from BIZON corresponds to about US$ 62.091 per GPU only if the total server cost is divided equally by eight—a comparative calculation, not the purchase price of each accelerator.
In summary: the references consulted indicate US$ 496.728 as the starting price of the X9000 G5 and US$ 544.280 as the median of HGX B300 node listings. In the index survey, the range observed was from US$ 399.999 to US$ 682.000 per node — about US$ 50 thousand to US$ 85.250 per GPU only as a system comparison baseline. For a purchase decision, request a formal quote with the component list, support modality, lead time, delivery, and applicable tax costs for the destination country.
Other BIZON servers for AI and LLMs
The X9000 G5 is the top-tier option based on HGX B300, but it's not BIZON's only AI line. In the catalog consulted in 11 of September of 2026, There are HGX servers with previous generations and configurable PCIe lines. The table below summarizes the families; availability, exact GPU, RAM, CPU, and network need to be confirmed in the quote.
| BIZON Line | Form factor and accelerators listed | Where it fits |
|---|---|---|
| X9000 G5 | HGX B300 with 8 B300 GPUs | Data center node for very large models, shared inference, training, and HPC. |
| X9000 G4 | HGX B200 with 8 B200 GPUs | Previous Blackwell alternative, also with NVLink/NVSwitch for distributed workloads across eight GPUs. |
| X9000 G3 | HGX H200 or H100 with 8 GPUs | Hopper platform for LLMs, training, and inference in data centers. |
| G9000 | Configurations of 4 or 8 GPUs A100, H100 or H200 | Configurable data center GPU server for AI, deep learning and parallel computing. |
| X7000 (G3) / G7000 G5 | Up to 8 PCIe GPUs; the X7000 is listed with RTX PRO 6000 Blackwell, A100, H100 and H200. The G7000 G5 lists RTX PRO 4000/4500/5000/6000 Blackwell, A100, H100 NVL and H200 NVL. | Training, inference and LLM services when PCIe GPU flexibility is more important than an HGX base. |
| X8000 / G8000 | Up to 4 PCIe GPUs; pages list RTX PRO 6000 Blackwell, A100, H100 and H200 | Inference, fine-tuning and teams that don't need eight accelerators in the same node. |
| Z9000 / ZX9000 | Server lines with liquid cooling and configurable NVIDIA GPUs | Environments where noise, temperature, and sustained workloads influence the system design. |
For smaller deployments or local LLMs, BIZON also advertises multi-GPU workstations such as the ZX5500. However, GPU count alone is not a sound basis for comparison: B300/HGX uses a data center interconnect that differs from PCIe cards. Memory per GPU, interconnect, model context, target latency, power, and support should guide the selection.
According to BIZON, BizonOS is based on Ubuntu 24.04 and includes components such as CUDA, Docker, PyTorch, TensorFlow, vLLM, Ollama, Hugging Face Transformers, Jupyter, and RAPIDS. This is a vendor software offering; before purchasing, confirm which versions, container images, and support levels are included in the selected configuration.
Infrastructure: the server is only part of the project
An 8U node with eight data center accelerators requires planning for rack space, power, cooling, noise, networking, and operations. BIZON lists twelve redundant 3.000 W power supplies and an operating environment of 10 °C to 30 °C; the power supply capacity is not a measurement of the system's power consumption. On its own page, the manufacturer gives an example range of 2.700 to 2.900 W for a specific configuration with eight B300 GPUs and two EPYC 9335 processors, at 100% load.
Before purchasing or deploying, validate the following with the vendor, an electrician, and the data center team: electrical circuits and redundancy, rack type, cooling capacity, airflow, low-latency networking, storage, firmware, drivers, observability, and the maintenance plan. This goes far beyond a workstation or local AI setup with a single GPU — for that different use case, compare it with our guide to refurbished AI server with 64 GB of VRAM.
Where a machine like this makes sense
- Shared inference: serving a large model to many concurrent users, APIs, or agents.
- Enterprise RAG: hosting a large language model within the company's environment, provided that the architecture includes access and data controls.
- Training and fine-tuning: workflows that benefit from large aggregate memory and fast communication between GPUs.
- HPC and research: accelerated simulations and pipelines that fit within the CUDA ecosystem.
To operate this type of service, monitor both the application and the infrastructure: queues, latency, request failures, GPU utilization, temperature, power, and networking. The guide to OpenObserve for logs, metrics, and traces is a starting point for planning telemetry, and the article on Prometheus and Node Exporter covers the Linux host foundation.
Technical checklist before evaluating a proposal
- Which model, version, context, quantization, and latency SLA must be supported?
- Does the test report TTFT, P95/P99, prompts, error rates, and the inference server version?
- Is the configuration AMD or Intel? How much RAM, NVMe storage, and networking is included?
- Can the rack, power, cooling, and acoustic conditions support the deployment?
- How will the GPUs, host, network, API, and queue be monitored? What is the incident response plan?
- What are the terms for support, replacement parts, firmware, and drivers?
Conclusion
The BIZON X9000 G5 shows what the HGX B300 class was designed to do: bring together eight Blackwell Ultra GPUs, ample HBM3e memory, and a fast interconnect in a single node for AI at scale. The video is useful because it provides concurrency figures, but they should be interpreted as the results of a specific demonstration. The real decision-making starts after the video: reproducing the actual workload, measuring latency as well as throughput, and designing the entire infrastructure around the server.
Sources
- BIZON — NVIDIA B300 Server Review: 8 GPUs, 128 Users, AI Benchmarks
- BIZON — X9000 G5
- NVIDIA — HGX AI Factory: HGX B300 components and specifications
- NVIDIA Developer Blog — Blackwell Ultra
- NVIDIA — Introduction to DGX B300
- Hashrate Index — AI Hardware Price Index (asking prices for HGX B300 nodes)
- BIZON — NVIDIA servers for AI, training, inference, and LLMs
- BIZON — BizonOS and the Ubuntu 24.04-based AI stack