Senior SRE, AIOps Platform for GPU Data Centers
nvidiaSanta Clara (CA)
Senior SRE, AIOps Platform for GPU Data Centers
Posted yesterday
nvidiaSanta Clara (CA)
Computer Systems Engineers/ArchitectsComputing Infrastructure Providers, Data Processing, Web Hosting, and Related Services
SENIORITY
Lead
About the role
NVIDIA Corporation in Santa Clara, CA is hiring a Dev Ops Engineer to operate our AI Data Center telemetry platform. You’ll own reliability, incident response, and postmortems for telemetry ingestion, processing, storage, and APIs/dashboards used by operators.
Expect to lead Kubernetes deployments end-to-end, build runbooks, and partner with Software and Systems Engineering to translate platform signals into actionable, trustworthy alerts and automation.
Before you apply
Applying takes about a minute. These four things decide how fast it moves after that.
Your profile is current
It's what we read first. Occupations, seniority and locations matter more than a long history.
Two examples you can talk through
Not a portfolio — just two pieces of work where you can explain the decisions and what you'd change.
A number in mind
What you're on now and what would make you move. We negotiate better when we know both.
Your notice period
Employers plan around it, and it's the question that stalls offers most often.
Once you apply, someone reads it and calls you before anything reaches the employer — usually within two working days.
More like this
