Senior Site Reliability Engineer (AWS)

Posted yesterday

meridianlinkDenver (CO)

SENIORITY

Senior

Apply

About the role

Senior Site Reliability Engineer As a Senior Site Reliability Engineer on our cloud engineering team, you'll keep our production environment healthy, secure, and running smoothly. This is an operations-focused role: you'll own the day-to-day administration of our AWS accounts and databases, backup posture across our data stores, and production monitoring and debugging for a fully serverless platform. Your work will span the operational side of the software development life cycle — from deployment to maintenance and updates — always striving for continuous improvement. You'll keep our infrastructure clean, easily deployable, and scalable, creating a stable operating environment for the whole team.
Responsibilities: Own day-to-day administration across AWS services, accounts, and access, as well as database administration across PostgreSQL and our other data stores. Own backup posture across databases, S3 buckets, and queues; verify restores regularly and maintain a tested disaster recovery plan. Proactively monitor production — Cloud Watch dashboards, metric alarms, log-based metrics, and Slack alerting — addressing operational issues before they impact users. Lead production debugging and incident response: build and maintain runbooks, participate in the on-call rotation, and resolve queue and dead-letter-queue failures through retry, redrive, and recovery. Continuously refine our infrastructure to ensure it is easily deployable and scalable: keep infrastructure as code (SST/Pulumi) accurate, retire unused infrastructure, and keep cost visible and justified. Share your knowledge of production operations with the team, fostering a culture of learning and growth.
Qualifications: Knowledge, Skills, & Abilities Bachelor's degree and 4-6 years of related experience or equivalent work experience.5+ years of experience in Dev Ops, site reliability, or platform operations, with significant responsibility for production systems.3+ years of hands-on experience with AWS, with an emphasis on serverless services (Lambda, SQS, Event Bridge, Cloud Watch, S3).Strong database administration experience: PostgreSQL operations, backup and recovery, and query performance; comfort administering other data stores. Proficiency in scripting languages such as Type Script, Python, and bash for production automation and operational tooling. Strong understanding of Linux, DNS, TLS, Docker, Git Hub Actions, and infrastructure as code (SST, Pulumi, or Terraform).Experience with production monitoring and alerting, incident response, and on-call ownership.

Before you apply

Applying takes about a minute. These four things decide how fast it moves after that.

Your profile is current

It's what we read first. Occupations, seniority and locations matter more than a long history.

Two examples you can talk through

Not a portfolio — just two pieces of work where you can explain the decisions and what you'd change.

A number in mind

What you're on now and what would make you move. We negotiate better when we know both.

Your notice period

Employers plan around it, and it's the question that stalls offers most often.

Once you apply, someone reads it and calls you before anything reaches the employer — usually within two working days.

More like this