Senior SRE, AIOps Platform for GPU Data Centers

Posted yesterday

nvidiaSanta Clara (CA)

SENIORITY

Lead

Apply

About the role

NVIDIA Corporation in Santa Clara, CA is hiring a Dev Ops Engineer to operate our AI Data Center telemetry platform. You’ll own reliability, incident response, and postmortems for telemetry ingestion, processing, storage, and APIs/dashboards used by operators. Expect to lead Kubernetes deployments end-to-end, build runbooks, and partner with Software and Systems Engineering to translate platform signals into actionable, trustworthy alerts and automation.

Before you apply

Applying takes about a minute. These four things decide how fast it moves after that.

Your profile is current

It's what we read first. Occupations, seniority and locations matter more than a long history.

Two examples you can talk through

Not a portfolio — just two pieces of work where you can explain the decisions and what you'd change.

A number in mind

What you're on now and what would make you move. We negotiate better when we know both.

Your notice period

Employers plan around it, and it's the question that stalls offers most often.

Once you apply, someone reads it and calls you before anything reaches the employer — usually within two working days.

More like this