After years of rapid advancement, AI has entered a new phase of maturity, moving from experimentation to real-world deployment. To help enterprises harness AI’s transformative potential, QCT offers its QCT AI POD, a software-defined, workload-driven infrastructure solution designed for production-ready AI environments.
The transformative power of AI across industries
Across manufacturing, healthcare, research and education, logistics, and retail, AI is automating workflows, improving efficiency, and accelerating innovation. Smart factories use agentic AI to optimize production, researchers leverage AI to solve complex challenges, and retailers enhance customer service with generative AI.
These technological breakthroughs are supported by high-performance AI infrastructures integrating some of the world’s most advanced computing, networking and storage technologies. However, building such powerful infrastructure on enterprise premises is highly challenging, as it requires deep hardware/software knowledge and involves rigorous testing and validation.
Accelerate your adoption journey with QCT AI POD
QCT AI POD is designed to streamline this process for enterprise customers. Leveraging years of hardware-software integration expertise, QCT develops its AI POD as a reference architecture for enterprises looking to adopt AI on their own trusted, workload-optimized infrastructure. It integrates compute, storage, networking, and cooling facilities into a cluster-level system design, all pre-configured to support a wide range of AI workloads. Furthermore, it combines pre-optimized hardware and software with system monitoring, management, and deployment tools, along with a robust AI development environment, helping enterprises reduce complexity and accelerate AI adoption.

Fig. 1. QCT AI POD overview
Throughout the development and deployment processes, QCT collaborates with ecosystem partners to offer end-to-end solutions, helping to shorten the journey from months to days. From initial demand pattern analysis and system design to system development, deployment, fine-tuning & performance benchmarking, and the final pilot run and go, QCT experts will provide reliable assistance, helping customers unlock greater agility, performance, and operational efficiency.
Choose the right AI POD configuration for your business
AI applications encompass a wide range of scenarios and require different hardware-software combinations to deliver the desired performance with power and cost efficiencies. To meet varying levels of workload requirements, QCT AI POD offers 4, 8, and 72-GPU rack-level configurations to handle medium, high, and extremely high compute demands. These configurations can be further customized to meet each customer’s unique requirements.

Fig. 2. QCT AI POD optimized for different scales
- 4-GPU AI POD:
Featuring two QuantaGrid D75E-4U as compute nodes, each equipped with four PCIe GPUs, this SKU supports a wide range of PCIe GPUs including NVIDIA RTX PRO 6000 Blackwell Series GPUs, NVIDIA H200 GPUs, and Intel® Gaudi® 3 AI accelerators. Paired with a utility node, a storage node, and an ethernet switch, this system powers computer vision, 3D visualization, and converged HPC and AI workloads with cost-effectiveness.

Fig. 3. QuantaGrid D75E-4U supporting multiple PCIe GPUs
- 8-GPU AI POD:
This category houses three SKUs. A PCIe GPU SKU equipped with four QuantaGrid D75E-4U is the most cost-effective option, optimized for retrieval-augmented generation (RAG), agentic AI, and LLM inference. For compute and memory-intensive workloads such as large language model (LLM) fine-tuning and deep learning training, 8 SXM GPU configurations are recommended, as the eight GPUs in each node are interconnected by NVIDIA NVLink to scale memory and performance. Customers can choose between the air-cooled or liquid-cooled configurations depending on their facility planning. The air-cooled version comprises three to four QuantaGrid D75H-10U per rack, while its liquid-cooled counterpart accommodates up to nine QuantaGrid D75L-3U (accelerated by NVIDIA HGX™ B300) or QuantaGrid D76T-2U (accelerated by NVIDIA HGX™ Rubin NVL8) per rack. For cooling facilities, QCT offers both in-rack coolant distribution units (CDUs) and its in-house developed QoolRack Sidecar to deliver efficient liquid-to-air cooling with smart power management.

Fig. 4. QuantaGrid D75H-10U for the air-cooled SKU

Fig. 5. QuantaGrid D75L-3U for the liquid-cooled SKU

Fig. 6. QuantaGrid D76T-2U for the liquid-cooled SKU
- 72-GPU AI POD:
Aligned with the NVIDIA NVL72 reference architecture, QCT’s 72-GPU AI POD incorporates 18 QuantaGrid D76V-1U accelerated by the NVIDIA Vera Rubin platform, or 18 QuantaGrid D75U-1U accelerated by the NVIDIA GB300 platform. It delivers the performance required for trillion-parameter LLM training, AI factory, agentic AI and long-context AI reasoning workloads. This liquid-cooled solution combines massive compute density, vast high-bandwidth memory (HBM) capacity, and advanced InfiniBand and Ethernet networking to power the most demanding AI applications. Its high-availability, in-band integrated switch design further ensures uninterrupted operation and enables large-scale, low-latency token generation at production scale.

Fig. 7. QuantaGrid D76V-1U accelerated by the latest NVIDIA Vera Rubin platform

Fig. 8. QuantaGrid D75U-1U accelerated by the NVIDIA GB300 platform
In addition to the highlighted features and components, QCT AI POD leverages open technologies to deliver comprehensive software services. It supports simplified cluster deployment, management, and real-time system monitoring, while providing ready-to-use toolkits and multiple user interface options to boost development efficiency.
Visit QCT’s AI POD solution page to kick start your AI transformation journey: QCT AI POD | QCT

