Senior Site Reliability Engineer

Posted yesterday

simarn solutionsSan Jose (CA)
Computer Systems Engineers/ArchitectsComputer Systems Design Services

SENIORITY

Lead

Apply

About the role

Job Title: Senior Site Reliability Engineer (SRE) | AI Infrastructure & KubernetesLocation: San Jose, CA ( Onsite)We're hiring an experienced SRE to build and scale reliable AI platforms on GCP. If you're passionate about automation, Kubernetes, cloud infrastructure, and observability, we'd love to hear from you.Key ResponsibilitiesAutomate infrastructure provisioning and deployments using Terraform and Ansible.Build and maintain monitoring and observability solutions.Design, deploy, and optimize AI workloads on Kubernetes.Improve scalability, performance, and reliability of AI services.Ensure security and compliance standards are met.Participate in SRE on-call rotations and incident response.Required SkillsStrong experience with Terraform, Ansible, and Infrastructure as Code (IaC).Hands-on expertise with GCP, Kubernetes, and Docker.Proficiency in Python and/or Go.Strong networking knowledge, including VPC, Shared VPC, and PSA.Understanding of cloud security, databases, and distributed systems.

Before you apply

Applying takes about a minute. These four things decide how fast it moves after that.

Your profile is current

It's what we read first. Occupations, seniority and locations matter more than a long history.

Two examples you can talk through

Not a portfolio — just two pieces of work where you can explain the decisions and what you'd change.

A number in mind

What you're on now and what would make you move. We negotiate better when we know both.

Your notice period

Employers plan around it, and it's the question that stalls offers most often.

Once you apply, someone reads it and calls you before anything reaches the employer — usually within two working days.

More like this