AI Data Center Manager
✓ Drop your CV once, then continue to the employer's application form. Your profile stays here for every recruiter hiring on igamingjobs.
Role in brief
BetConstruct is seeking an AI Data Center Manager to oversee a high-density AI infrastructure project. Candidates with extensive experience in data center operations and high-performance computing should apply.
About the role
In this role, you will manage a 5MW high-density compute environment focused on enterprise AI workloads, utilizing advanced technologies such as liquid cooling and Kubernetes. Your leadership will be crucial in optimizing power distribution and achieving sustainability goals.
You will work closely with network architecture teams and colocation partners to ensure the seamless integration of high-performance computing systems. Success in this position will be measured by your ability to maintain operational efficiency and uphold infrastructure uptime.
Skills that matter here
- AI: This role requires a strong understanding of AI workloads to effectively manage the associated infrastructure.
- data center operations: You will oversee the daily operations and maintenance of a high-density data center.
- high-performance computing: Experience in HPC is essential for managing the specialized compute environments.
- liquid cooling: You will implement and maintain advanced liquid cooling systems for optimal performance.
- Kubernetes: Knowledge of Kubernetes is necessary for managing container orchestration in the data center.
- Slurm: You will utilize Slurm for scheduling AI training tasks and managing compute resources.
Who this role suits
- A seasoned leader with over 7 years of experience in data center operations or HPC.
- Someone with a strong technical background in liquid cooling and power distribution systems.
- A proactive problem solver who thrives in high-pressure environments and values sustainability.
- An individual with established relationships in the industry, particularly with component providers and server OEMs.
From the employer
Responsibilities
- Oversee the operation and maintenance of direct-to-chip liquid cooling loops, Coolant Distribution Units (CDUs), and localized liquid-to-air secondary heat exchangers in a dense environment.
- Optimize active power distribution at the rack level, managing ultra-dense power envelopes ranging from 60kW to 100kW+ per rack across our 5MW footprint.
- Drive continuous physical optimization to lower Power Usage Effectiveness (PUE) and achieve carbon neutrality and environmental sustainability metrics.
- Act as primary liaison to wholesale colocation partners, holding suppliers strictly accountable to power, cooling, physical security, and infrastructure uptime SLAs.
- Direct the logistics, installation, staging, provisioning, and decommissioning of advanced AI hardware systems (including NVIDIA DGX/Blackwell/Hopper, Ver Rubin).
- Collaborate closely with Network Architecture teams to ensure high-bandwidth, ultra-low-latency backend topologies (InfiniBand, RoCEv2) are perfectly integrated and structurally protected.
- Implement and maintain centralized environmental telemetry pipelines that connect physical variables such as coolant flow rates, pressure, and ambient temperature with software schedulers such as Slurm and Kubernetes.
Requirements
- 7+ years of direct experience in data center operations, high-performance computing (HPC), or critical facilities engineering, with at least 3 years in a direct leadership or managerial capacity.
- Deep technical understanding of direct-to-chip liquid cooling loops, Coolant Distribution Units, rear-door heat exchangers (RDHx), and general facility MEP (Mechanical, Electrical, Plumbing) designs.
- Direct experience deploying and maintaining high-power density compute environments (NVIDIA HGX/DGX, advanced liquid-cooled chassis, or specialized OEM architectures).
- Proven experience negotiating and managing SLAs, Master Service Agreements (MSAs), and joint operating procedures inside high-tier commercial colocation facilities.
- Bachelor's degree in Mechanical Engineering, Electrical Engineering, Computer Science, or equivalent technical discipline / equivalent practical experience.
- Experience guiding active environments through transitioning from legacy air cooling to hybrid air-and-liquid or 100% direct-to-chip liquid-cooled architectures.
- Functional knowledge of container orchestrators (Kubernetes) and AI training scheduling architectures (Slurm).
- Established, active relationships with key component providers (CDU vendors, quick-disconnect suppliers) and server OEMs.
Questions about this role
What is the remote policy?
This position is fully remote.
What level of experience is required?
Candidates should have at least 7 years of experience in data center operations or HPC, with 3 years in a managerial role.
How do I apply?
You can apply through the BetConstruct careers page linked in the job posting.
✓ Drop your CV once, then continue to the employer's application form. Your profile stays here for every recruiter hiring on igamingjobs.