{"id":1564,"date":"2026-09-11T19:35:31","date_gmt":"2026-09-11T22:35:31","guid":{"rendered":"https:\/\/www.linuxpro.com.br\/?p=1564"},"modified":"2026-09-11T19:48:08","modified_gmt":"2026-09-11T22:48:08","slug":"nvidia-b300-o-servidor-bizon-x9000-g5-com-8-gpus-para-ia","status":"publish","type":"post","link":"https:\/\/www.linuxpro.com.br\/en\/2026\/09\/nvidia-b300-o-servidor-bizon-x9000-g5-com-8-gpus-para-ia\/","title":{"rendered":"NVIDIA B300: the BIZON X9000 G5 server with 8 GPUs for AI"},"content":{"rendered":"<p><img decoding=\"async\" src=\"\/wp-content\/uploads\/2026\/09\/b300-bizon-x9000-g5-v2.webp\" alt=\"Mascote LinuxPro ao lado de um servidor de IA com oito GPUs e um cachorro caramelo em um data center\" \/><\/p>\n<p>A server with eight NVIDIA B300 GPUs is not a \u201cgraphics card machine\u201d: it is data center infrastructure for serving very large AI models to many users simultaneously. In this video, BIZON shows the <strong>X9000 G5<\/strong>, with eight NVIDIA B300 GPUs on an <strong>HGX B300<\/strong>, platform, running Kimi K3 with vLLM and measuring throughput for 1 to 128 concurrent users. The demo is interesting, but also a good opportunity to separate architecture, commercial server, and marketing benchmark.<\/p>\n<blockquote>\n<p><strong>Video summary:<\/strong> \u201cNVIDIA B300 Server Review: 8 GPUs, 128 Users, AI Benchmarks\u201d, from the BIZON channel, published on 7 of September 2026. The manufacturer presents the BIZON X9000 G5 with eight NVIDIA B300 GPUs on an HGX B300 platform and inference results with Kimi K3 and vLLM. <a href=\"https:\/\/www.youtube.com\/watch?v=5rGBUCFJJrk\" target=\"_blank\" rel=\"noopener\">Watch on YouTube<\/a>.<\/p>\n<\/blockquote>\n<div class=\"video-container\" style=\"position:relative;padding-bottom:56.25%;height:0;overflow:hidden;margin:1.5em 0;\"><iframe src=\"https:\/\/www.youtube.com\/embed\/5rGBUCFJJrk\" title=\"NVIDIA B300 Server Review: 8 GPUs, 128 Users, AI Benchmarks\" loading=\"lazy\" frameborder=\"0\" allow=\"accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share\" referrerpolicy=\"strict-origin-when-cross-origin\" allowfullscreen style=\"position:absolute;top:0;left:0;width:100%;height:100%;\"><\/iframe><\/div>\n<h2>What are NVIDIA B300, HGX B300 and BIZON X9000 G5?<\/h2>\n<p>The names sound similar, but they describe different layers of the solution:<\/p>\n<ul>\n<li><strong>NVIDIA B300:<\/strong> GPU from the Blackwell Ultra generation aimed at AI, HPC, and data center inference.<\/li>\n<li><strong>NVIDIA HGX B300:<\/strong> platform\/baseboard with eight B300 GPUs connected via NVLink and NVSwitch.<\/li>\n<li><strong>BIZON X9000 G5:<\/strong> BIZON server that integrates the HGX B300 platform in its own chassis.<\/li>\n<li><strong>NVIDIA DGX B300:<\/strong> another product: a complete NVIDIA system. It is not the X9000 G5 and should not be treated as if it were the same machine.<\/li>\n<\/ul>\n<p>This distinction matters when comparing CPU, system memory, network, power supplies, support, and software. Two machines can use HGX B300 but remain different products with different integration and support choices.<\/p>\n<h2>What the HGX B300 platform delivers<\/h2>\n<p>In NVIDIA's reference architecture, an HGX B300 node packs eight Blackwell Ultra GPUs. Each GPU has <strong>288 GB of HBM3e<\/strong>, reaching <strong>2,30 TB of HBM3e memory on the node<\/strong>. The documentation also indicates up to 8 TB\/s of memory bandwidth per GPU and 64 TB\/s in aggregate across the GPU set.<\/p>\n<p>For large models, memory capacity is not enough: GPUs need to exchange activations, partitioned weights, and communication data without turning the external network into a bottleneck. That's why HGX B300 uses fifth-generation NVLink and NVSwitch. NVIDIA specifies 1,8 TB\/s of GPU-to-GPU communication and 14,4 TB\/s aggregate across the eight-GPU domain. To scale beyond a single node, the reference includes eight ConnectX-8 SuperNICs, with up to 800 Gb\/s per adapter.<\/p>\n<p>It's the difference between placing many independent GPUs in a server and building an accelerated domain that can cooperate when serving an LLM. To understand why memory, interconnection, and numeric format matter so much in AI, see also <a href=\"\/en\/2026\/09\/gpu-vs-tpu-como-funcionam-chips-ia\/\">GPU vs TPU: how AI chips work<\/a>.<\/p>\n<h2>What the BIZON X9000 G5 adds to the mix<\/h2>\n<p>The <a href=\"https:\/\/bizon-tech.com\/bizon-x9000-g5.html\" target=\"_blank\" rel=\"noopener\">X9000 G5<\/a> is the BIZON integration in 8U form factor. In the AMD configuration listed by the manufacturer, there is support for two EPYC 9005\/9004 processors, up to 3 TB of DDR5 ECC, eight front NVMe Gen5 hot-swap bays and four PCIe Gen5 x16 slots. The page also lists network options with eight OSFP ports of 800 Gb\/s for InfiniBand XDR or two Ethernet links of 400 Gb\/s per ConnectX-8.<\/p>\n<p>There is also an Intel variant listed by BIZON, with two Xeon 6500\/6700 and up to 4 TB of DDR5. Therefore, it makes no sense to talk about \u201cthe\u201d amount of CPU RAM of the X9000 G5 without saying which configuration is being quoted. Items such as CPU, RAM, storage, adapters, and support contract are part of the project, not footnote details.<\/p>\n<h2>The video benchmark: Kimi K3 with vLLM<\/h2>\n<p>BIZON uses the <a href=\"\/en\/2026\/09\/vllm-primeiros-passos-linux\/\">vLLM<\/a> as an inference server and claims to load the Kimi K3 on the eight GPU cluster. The central metric disclosed is <strong>aggregate throughput<\/strong>, in tokens per second, as the number of concurrent users increases:<\/p>\n<table>\n<thead>\n<tr>\n<th>Concurrent users<\/th>\n<th>Disclosed aggregate throughput<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>1<\/td>\n<td>99,8 tokens\/s<\/td>\n<\/tr>\n<tr>\n<td>32<\/td>\n<td>2.058,7 tokens\/s<\/td>\n<\/tr>\n<tr>\n<td>64<\/td>\n<td>3.013,4 tokens\/s<\/td>\n<\/tr>\n<tr>\n<td>128<\/td>\n<td>4.508,8 tokens\/s<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>These values are <strong>results reported by the manufacturer in the video<\/strong> for a single workload; they are not 4.508,8 tokens\/s per person and do not automatically measure the experience of any application. Even so, they illustrate an important characteristic of inference servers: under concurrency, continuous batch scheduling can greatly increase the total tokens produced per second.<\/p>\n<h2>Why tokens per second is not enough<\/h2>\n<p>A throughput table is only part of the test. To decide whether a platform serves a chatbot, an agent API or an internal team, you also need to measure <em>time to first token<\/em> (TTFT), per-token latency, P95\/P99 percentiles, error rate, context length, input and output size, queue policy, and actual consumption.<\/p>\n<p>The video does not publish enough details for independent reproduction: exact model version and quantization, vLLM version and parameters, prompts, output limit, concurrency method, TTFT and latency percentiles are not documented in the description. Therefore, the correct reading is: <strong>a capability demonstration of the BIZON + HGX B300 + chosen workload<\/strong>, and not a universal ranking between GPUs or a capacity promise for all models.<\/p>\n<h2>How much does the BIZON X9000 G5 and each B300 cost?<\/h2>\n<p>Corporate AI hardware pricing is not very transparent: there is no public retail price list for B300 in the references consulted, and OEMs and integrators sell these systems through quotes. Therefore, the correct value for planning is a <strong>reference range<\/strong>, followed by a quote for the exact configuration.<\/p>\n<p>In the query made on <strong>11 of September of 2026<\/strong>, the BIZON page showed the <a href=\"https:\/\/bizon-tech.com\/bizon-x9000-g5.html\" target=\"_blank\" rel=\"noopener\">X9000 G5 starting at US$ 496.728<\/a>. It's the starting price of a complete server, not of eight standalone GPUs: CPU, system memory, NVMe, network, chassis, power supplies, integration and support also make up the proposal. Freight, taxes, import and local services may change the final cost.<\/p>\n<p>As a market reference, the <a href=\"https:\/\/hashrateindex.com\/blog\/announcement-introducing-the-ai-hardware-price-index\/\" target=\"_blank\" rel=\"noopener\">AI Hardware Price Index<\/a> recorded in 31 August 2026 a median of <strong>US$ 544.280 per new HGX B300 node with eight GPUs<\/strong>, equivalent to <strong>US$ 68.035 per GPU<\/strong> in the index metric. In that survey, the advertisements ranged from US$ 399.999 to US$ 682.000 per node.<\/p>\n<p>This figure of US$ 68.035 <strong>is neither MSRP nor the price of a standalone B300<\/strong>: it's the division of the advertised price of complete eight-GPU systems. The index itself notes that it works with asking prices from sellers, not completed sales, and doesn't standardize CPU, RAM, storage, or warranty. The initial price of US$ 496.728 from BIZON corresponds to about US$ 62.091 per GPU only if the total server cost is divided equally by eight\u2014a comparative calculation, not the purchase price of each accelerator.<\/p>\n<p>In summary: the references consulted indicate <strong>US$ 496.728 as the starting price of the X9000 G5<\/strong> and <strong>US$ 544.280 as the median of HGX B300 node listings<\/strong>. In the index survey, the range observed was from <strong>US$ 399.999 to US$ 682.000 per node<\/strong> \u2014 about <strong>US$ 50 thousand to US$ 85.250 per GPU<\/strong> only as a system comparison baseline. For a purchase decision, request a formal quote with the component list, support modality, lead time, delivery, and applicable tax costs for the destination country.<\/p>\n<h2>Other BIZON servers for AI and LLMs<\/h2>\n<p>The X9000 G5 is the top-tier option based on HGX B300, but it's not BIZON's only AI line. In the catalog consulted in <strong>11 of September of 2026<\/strong>, There are HGX servers with previous generations and configurable PCIe lines. The table below summarizes the families; availability, exact GPU, RAM, CPU, and network need to be confirmed in the quote.<\/p>\n<table>\n<thead>\n<tr>\n<th>BIZON Line<\/th>\n<th>Form factor and accelerators listed<\/th>\n<th>Where it fits<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><a href=\"https:\/\/bizon-tech.com\/bizon-x9000-g5.html\" target=\"_blank\" rel=\"noopener\">X9000 G5<\/a><\/td>\n<td>HGX B300 with 8 B300 GPUs<\/td>\n<td>Data center node for very large models, shared inference, training, and HPC.<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/bizon-tech.com\/deep-learning-nvidia-gpu-servers\" target=\"_blank\" rel=\"noopener\">X9000 G4<\/a><\/td>\n<td>HGX B200 with 8 B200 GPUs<\/td>\n<td>Previous Blackwell alternative, also with NVLink\/NVSwitch for distributed workloads across eight GPUs.<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/bizon-tech.com\/bizon-x9000.html\" target=\"_blank\" rel=\"noopener\">X9000 G3<\/a><\/td>\n<td>HGX H200 or H100 with 8 GPUs<\/td>\n<td>Hopper platform for LLMs, training, and inference in data centers.<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/bizon-tech.com\/bizon-g9000.html\" target=\"_blank\" rel=\"noopener\">G9000<\/a><\/td>\n<td>Configurations of 4 or 8 GPUs A100, H100 or H200<\/td>\n<td>Configurable data center GPU server for AI, deep learning and parallel computing.<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/bizon-tech.com\/deep-learning-nvidia-gpu-servers\" target=\"_blank\" rel=\"noopener\">X7000 (G3)<\/a> \/ <a href=\"https:\/\/bizon-tech.com\/bizon-g7000-g5.html\" target=\"_blank\" rel=\"noopener\">G7000 G5<\/a><\/td>\n<td>Up to 8 PCIe GPUs; the X7000 is listed with RTX PRO 6000 Blackwell, A100, H100 and H200. The G7000 G5 lists RTX PRO 4000\/4500\/5000\/6000 Blackwell, A100, H100 NVL and H200 NVL.<\/td>\n<td>Training, inference and LLM services when PCIe GPU flexibility is more important than an HGX base.<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/bizon-tech.com\/bizon-x8000.html\" target=\"_blank\" rel=\"noopener\">X8000 \/ G8000<\/a><\/td>\n<td>Up to 4 PCIe GPUs; pages list RTX PRO 6000 Blackwell, A100, H100 and H200<\/td>\n<td>Inference, fine-tuning and teams that don't need eight accelerators in the same node.<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/bizon-tech.com\/liquid-cooled-workstation-computers\" target=\"_blank\" rel=\"noopener\">Z9000 \/ ZX9000<\/a><\/td>\n<td>Server lines with liquid cooling and configurable NVIDIA GPUs<\/td>\n<td>Environments where noise, temperature, and sustained workloads influence the system design.<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>For smaller deployments or local LLMs, BIZON also advertises multi-GPU workstations such as the ZX5500. However, GPU count alone is not a sound basis for comparison: B300\/HGX uses a data center interconnect that differs from PCIe cards. Memory per GPU, interconnect, model context, target latency, power, and support should guide the selection.<\/p>\n<p>According to BIZON, BizonOS is based on Ubuntu 24.04 and includes components such as CUDA, Docker, PyTorch, TensorFlow, vLLM, Ollama, Hugging Face Transformers, Jupyter, and RAPIDS. This is a vendor software offering; before purchasing, confirm which versions, container images, and support levels are included in the selected configuration.<\/p>\n<h2>Infrastructure: the server is only part of the project<\/h2>\n<p>An 8U node with eight data center accelerators requires planning for rack space, power, cooling, noise, networking, and operations. BIZON lists twelve redundant 3.000 W power supplies and an operating environment of 10 \u00b0C to 30 \u00b0C; the power supply capacity is not a measurement of the system's power consumption. On its own page, the manufacturer gives an example range of 2.700 to 2.900 W for a specific configuration with eight B300 GPUs and two EPYC 9335 processors, at 100% load.<\/p>\n<p>Before purchasing or deploying, validate the following with the vendor, an electrician, and the data center team: electrical circuits and redundancy, rack type, cooling capacity, airflow, low-latency networking, storage, firmware, drivers, observability, and the maintenance plan. This goes far beyond a workstation or local AI setup with a single GPU \u2014 for that different use case, compare it with our guide to <a href=\"\/en\/2026\/09\/servidor-ia-recondicionado-64gb-vram-v100-p100-mi25\/\">refurbished AI server with 64 GB of VRAM<\/a>.<\/p>\n<h2>Where a machine like this makes sense<\/h2>\n<ul>\n<li><strong>Shared inference:<\/strong> serving a large model to many concurrent users, APIs, or agents.<\/li>\n<li><strong>Enterprise RAG:<\/strong> hosting a large language model within the company's environment, provided that the architecture includes access and data controls.<\/li>\n<li><strong>Training and fine-tuning:<\/strong> workflows that benefit from large aggregate memory and fast communication between GPUs.<\/li>\n<li><strong>HPC and research:<\/strong> accelerated simulations and pipelines that fit within the CUDA ecosystem.<\/li>\n<\/ul>\n<p>To operate this type of service, monitor both the application and the infrastructure: queues, latency, request failures, GPU utilization, temperature, power, and networking. The guide to <a href=\"\/en\/2026\/09\/openobserve-observabilidade-logs-metricas-traces\/\">OpenObserve for logs, metrics, and traces<\/a> is a starting point for planning telemetry, and the article on <a href=\"\/en\/2026\/09\/monitorando-servidores-linux-com-prometheus\/\">Prometheus and Node Exporter<\/a> covers the Linux host foundation.<\/p>\n<h2>Technical checklist before evaluating a proposal<\/h2>\n<ul>\n<li>Which model, version, context, quantization, and latency SLA must be supported?<\/li>\n<li>Does the test report TTFT, P95\/P99, prompts, error rates, and the inference server version?<\/li>\n<li>Is the configuration AMD or Intel? How much RAM, NVMe storage, and networking is included?<\/li>\n<li>Can the rack, power, cooling, and acoustic conditions support the deployment?<\/li>\n<li>How will the GPUs, host, network, API, and queue be monitored? What is the incident response plan?<\/li>\n<li>What are the terms for support, replacement parts, firmware, and drivers?<\/li>\n<\/ul>\n<h2>Conclusion<\/h2>\n<p>The BIZON X9000 G5 shows what the HGX B300 class was designed to do: bring together eight Blackwell Ultra GPUs, ample HBM3e memory, and a fast interconnect in a single node for AI at scale. The video is useful because it provides concurrency figures, but they should be interpreted as the results of a specific demonstration. The real decision-making starts after the video: reproducing the actual workload, measuring latency as well as throughput, and designing the entire infrastructure around the server.<\/p>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/www.youtube.com\/watch?v=5rGBUCFJJrk\" target=\"_blank\" rel=\"noopener\">BIZON \u2014 NVIDIA B300 Server Review: 8 GPUs, 128 Users, AI Benchmarks<\/a><\/li>\n<li><a href=\"https:\/\/bizon-tech.com\/bizon-x9000-g5.html\" target=\"_blank\" rel=\"noopener\">BIZON \u2014 X9000 G5<\/a><\/li>\n<li><a href=\"https:\/\/docs.nvidia.com\/enterprise-reference-architectures\/hgx-ai-factory\/latest\/components.html\" target=\"_blank\" rel=\"noopener\">NVIDIA \u2014 HGX AI Factory: HGX B300 components and specifications<\/a><\/li>\n<li><a href=\"https:\/\/developer.nvidia.com\/blog\/inside-nvidia-blackwell-ultra-the-chip-powering-the-ai-factory-era\/\" target=\"_blank\" rel=\"noopener\">NVIDIA Developer Blog \u2014 Blackwell Ultra<\/a><\/li>\n<li><a href=\"https:\/\/docs.nvidia.com\/dgx\/dgxb300-user-guide\/introduction-to-dgxb300.html\" target=\"_blank\" rel=\"noopener\">NVIDIA \u2014 Introduction to DGX B300<\/a><\/li>\n<li><a href=\"https:\/\/hashrateindex.com\/blog\/announcement-introducing-the-ai-hardware-price-index\/\" target=\"_blank\" rel=\"noopener\">Hashrate Index \u2014 AI Hardware Price Index (asking prices for HGX B300 nodes)<\/a><\/li>\n<li><a href=\"https:\/\/bizon-tech.com\/deep-learning-nvidia-gpu-servers\" target=\"_blank\" rel=\"noopener\">BIZON \u2014 NVIDIA servers for AI, training, inference, and LLMs<\/a><\/li>\n<li><a href=\"https:\/\/bizon-tech.com\/about\" target=\"_blank\" rel=\"noopener\">BIZON \u2014 BizonOS and the Ubuntu 24.04-based AI stack<\/a><\/li>\n<\/ul>","protected":false},"excerpt":{"rendered":"<p>Understand the NVIDIA B300, the HGX B300 platform, and the BIZON X9000 G5. We analyze the video featuring a Kimi K3 and vLLM benchmark for up to 128 users, along with memory, networking, power, and evaluation criteria.<\/p>","protected":false},"author":1,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[115,21,120],"tags":[134,116,385,128,152,384],"class_list":["post-1564","post","type-post","status-publish","format-standard","hentry","category-ia","category-infra","category-servidores","tag-gpu","tag-ia","tag-inferencia","tag-llm","tag-nvidia","tag-vllm"],"_links":{"self":[{"href":"https:\/\/www.linuxpro.com.br\/en\/wp-json\/wp\/v2\/posts\/1564","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.linuxpro.com.br\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.linuxpro.com.br\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.linuxpro.com.br\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.linuxpro.com.br\/en\/wp-json\/wp\/v2\/comments?post=1564"}],"version-history":[{"count":7,"href":"https:\/\/www.linuxpro.com.br\/en\/wp-json\/wp\/v2\/posts\/1564\/revisions"}],"predecessor-version":[{"id":1573,"href":"https:\/\/www.linuxpro.com.br\/en\/wp-json\/wp\/v2\/posts\/1564\/revisions\/1573"}],"wp:attachment":[{"href":"https:\/\/www.linuxpro.com.br\/en\/wp-json\/wp\/v2\/media?parent=1564"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.linuxpro.com.br\/en\/wp-json\/wp\/v2\/categories?post=1564"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.linuxpro.com.br\/en\/wp-json\/wp\/v2\/tags?post=1564"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}