Technical Program Manager – AI Infrastructure / GPU Clusters
Technical Program Manager – AI Infrastructure / GPU Clusters
Posted 5 days ago
SENIORITY
Manager
About the role
- GPU Cluster Deployment Lead the end-to-end deployment of AI GPU clusters, from infrastructure planning through production launch.
- Drive coordination across Infrastructure Solution Architects, network engineers, hardware vendors, and data center teams.
- Manage delivery timelines covering hardware deployment, network integration, cluster bring-up, and production readiness.
- Infrastructure Architecture Collaboration Work closely with Infrastructure Solution Architects (SA) to define: GPU server platform selection
- Network architecture for distributed GPU clusters
- Storage integration and cluster infrastructure design
- Support development of the cluster Bill of Materials (BOM) including compute, networking, storage, and supporting infrastructure components.
- Ensure architecture decisions align with data center constraints such as power density, cooling capacity, and rack layout.
- System Integration Drive system integration for large-scale GPU clusters, including: Rack elevation planningGPU server deployment and configuration
- High-speed network topology implementation
- Power and cooling readiness
- Ensure deployments align with vendor reference architectures and validated cluster designs.
- Contractor & Field Deployment Management Work closely with General Contractors (GC) and system integrators to manage on-site infrastructure implementation.
- Lead contractor onboarding, including SOW development, scope definition, and delivery milestone alignment.
- Coordinate and oversee field deployment activities such as: Structured cabling installation
- Rack installation and equipment mounting
- Network and power connectivity preparation
- Hardware staging and deployment logistics
- Cluster Validation & Performance Testing Coordinate cluster bring-up and validation activities including: Single-node GPU validation
- Multi-node cluster deploymentGPU interconnect validation (P2P, RDMA)
- Drive cluster benchmarking, stress testing, and performance verification before production release.
- Operational Readiness Ensure deployed GPU clusters are fully ready for production workloads by driving: Hardware and network validation
- Monitoring and telemetry integration
- Operational documentation and runbooks
- Handover to operations teams Qualifications 5+ years experience in Technical Program Management, Infrastructure Program Management, or HPC infrastructure delivery
- Experience with GPU cluster deployments or high-performance computing environments
- Familiarity with GPU server architecture and distributed computing infrastructure
- Experience working with Infrastructure Solution Architects to define system architecture and hardware BOMExperience managing data center hardware deployments and system integration
- Ability to coordinate multi-vendor infrastructure projects across regions Required SkillsExperience deploying large-scale AI infrastructure or GPU clusters
- Familiarity with: InfiniBand / RoCE / high-speed Ethernet networkingGPU interconnect validation (P2P / RDMA)
- Rack elevation and high-density rack deployment
- Experience with cluster validation and performance benchmarking
- Background as Systems Engineer, HPC Engineer, or Infrastructure Architect
- Experience working in AI infrastructure, cloud infrastructure, or hyperscale data centers Preferred SkillsExperience deploying liquid-cooled GPU clusters or high-power racks
- Experience working with NVIDIA AI infrastructure platforms
- Familiarity with AI training environments and distributed workloads
Before you apply
Applying takes about a minute. These four things decide how fast it moves after that.
Your profile is current
It's what we read first. Occupations, seniority and locations matter more than a long history.
Two examples you can talk through
Not a portfolio — just two pieces of work where you can explain the decisions and what you'd change.
A number in mind
What you're on now and what would make you move. We negotiate better when we know both.
Your notice period
Employers plan around it, and it's the question that stalls offers most often.
Once you apply, someone reads it and calls you before anything reaches the employer — usually within two working days.
More like this
