Senior Site Reliability Engineer

Posted yesterday

goguardianEl Segundo (CA)

SENIORITY

Lead

Apply

About the role

We're looking for a Senior Site Reliability Engineer (SRE) to help design, scale, and maintain the infrastructure that powers our core products and services. In this role, you'll collaborate with engineering teams to drive operational excellence, optimise system performance, and ensure high availability across production environments. This position sits on Tech Foundation, a team that manages core cloud infrastructure, shared data services, and developer tooling to empower our product teams to deliver software efficiently and securely. The ideal candidate brings a strong background in cloud infrastructure, automation, and modern reliability practices, with a passion for solving complex operational challenges in a collaborative environment.
Responsibilities: Architect and maintain scalable, secure cloud infrastructure to ensure high availability for core products. Enhance observability and monitoring frameworks to deliver highly accurate alerts, minimising noise and improving incident detection. Participate in on-call rotations and lead incident response, ensuring comprehensive post-mortems and RCAs are completed to drive systemic improvements. Optimise and modernise deployment pipelines and automation workflows to maximise engineering velocity and operational safety. Partner with product development teams to provide infrastructure support, review architectural changes, and promote reliability best practices. Implement and uphold robust security standards and compliance controls across all managed cloud infrastructure.
Requirements: 5+ years of professional experience in Site Reliability Engineering, Infrastructure, or DevOps roles supporting production SaaS applications. Strong proficiency with AWS core services (including EC2 VPC, S3) along with experience in Serverless frameworks and managed Kubernetes environments like EKS. Extensive experience writing and managing Infrastructure as Code (IaC) using Terraform. Familiarity with configuring, troubleshooting, and maintaining data layers such as MongoDB, Redshift, and OpenSearch. Experience with GCP environments or technologies like Firestore is a plus. Experience managing or modernising CI/CD pipelines and deployment workflows utilising systems like Jenkins, AWS CodeBuild/CodePipeline, or GitHub Actions. Deep understanding of Linux operating system fundamentals and Unix shell scripting. Ability to read and debug code written in JavaScript/TypeScript, Python, or Go to effectively troubleshoot underlying service errors. Strong communication and collaboration skills, with a track record of driving technical decisions and establishing team-wide operational standards. Eager to take initiative in a fast-paced, ever-changing, dynamic environment. Fueled by the opportunity to truly impact the education landscape.

Before you apply

Applying takes about a minute. These four things decide how fast it moves after that.

Your profile is current

It's what we read first. Occupations, seniority and locations matter more than a long history.

Two examples you can talk through

Not a portfolio — just two pieces of work where you can explain the decisions and what you'd change.

A number in mind

What you're on now and what would make you move. We negotiate better when we know both.

Your notice period

Employers plan around it, and it's the question that stalls offers most often.

Once you apply, someone reads it and calls you before anything reaches the employer — usually within two working days.

More like this