In AI data centers, managing distributed GPU-powered machine learning frameworks is a central challenge. Data scientists run diverse workloads all at once. From data preparation, model training, model validation and inference, these workloads need to run quickly, use resources efficiently, and account for factors like CPU and GPU architecture, memory, cache, bus topologies and NVIDIA interconnect and network switch topologies.
NVIDIA DGX systems are purpose-built to meet these demands. Getting the most out of them requires equally capable workload management.
HPCWorks PBS Professional is a fast, powerful workload manager designed to improve productivity, optimize utilization and efficiency, and simplify administration. From the biggest HPC workloads to millions of small, high-throughput jobs.
It automates job scheduling, management, monitoring and reporting, and is the trusted solution for complex Top500 systems as well as smaller clusters.
Key capabilities include: