HPC Data Center Developer

Posted 3 days ago

autonomai recruitmentNew York (NY)

SENIORITY

Senior

Apply

About the role

HPC Data Center Developer | Global Trading Firm Chicago or New York | We’re partnering with a global trading firm to hire an HPC Data Center Developer to build the software and automation behind its large-scale compute infrastructure. This is a development-heavy role with ownership of production systems spanning hardware provisioning, infrastructure monitoring, capacity planning and failure simulation. You’ll write substantial amounts of code and work closely with HPC engineering and operations teams to solve problems across the physical infrastructure powering research and trading.
The role: Build end-to-end hardware provisioning workflows covering discovery, configuration, validation and deployment across servers, switches, power and cooling equipment. Develop tools to model power and cooling capacity, forecast demand and identify infrastructure constraints. Create failure simulation tooling to assess the impact of outages and validate infrastructure resilience. Integrate hardware telemetry and external data feeds into central monitoring platforms, building dashboards, exporters and alerting. Automate hardware lifecycle tracking, inventory management, diagnostics and operational workflows. Own the reliability, maintenance and ongoing development of the systems you build. Use AI development tools daily to support coding, debugging, analysis and documentation.
What we’re looking for: 5+ years in infrastructure software development, production engineering, automation or SRE, ideally within HPC or large-scale data center environments. A track record of shipping reliable, maintainable production tooling. Strong programming skills in Go and Python, or comparable infrastructure development experience. Deep Linux knowledge, including networking, storage, system administration and troubleshooting. Experience automating hardware provisioning and working with interfaces such as Redfish, IPMI/BMC, SNMP or vendor APIs. An understanding of data center power, cooling and physical infrastructure, including air and liquid cooling. Experience with observability platforms such as Grafana, Prometheus or InfluxDB, alongside configuration management and infrastructure-as-code tools. Strong networking fundamentals and confidence integrating APIs, databases and infrastructure metrics. Experience with ClickHouse, MySQL, Arista/Cisco networking and CI/CD workflows would also be valuable. You’ll suit this role if you enjoy understanding how infrastructure works from the hardware through to the software, investigating root causes and turning operational challenges into well-engineered systems. Participation in coordinated evening and weekend maintenance is required. Interested? Apply or message me directly for a confidential conversation.

Before you apply

Applying takes about a minute. These four things decide how fast it moves after that.

Your profile is current

It's what we read first. Occupations, seniority and locations matter more than a long history.

Two examples you can talk through

Not a portfolio — just two pieces of work where you can explain the decisions and what you'd change.

A number in mind

What you're on now and what would make you move. We negotiate better when we know both.

Your notice period

Employers plan around it, and it's the question that stalls offers most often.

Once you apply, someone reads it and calls you before anything reaches the employer — usually within two working days.

More like this