Engineer - Cluster Operations

Posted yesterday

tata consultancy servicesSeattle (WA)
Computer Systems Engineers/ArchitectsComputer Systems Design Services

SENIORITY

Senior

SALARY

$120,000-$160,000 a year

Apply

About the role

Monitoring & Troubleshooting (CloudWatch/Logs) Analyze EMR cluster metrics, Spark application telemetry, and YARN resource utilization to identify over-provisioned memory allocations, underutilized executors, and suboptimal cluster configurations that contribute to excessive DRAM consumption. Recommend and implement cluster-level optimizations, including instance family right-sizing (e.g., migrating from memory-optimized R-type to compute-optimized C-type instances), node count adjustments, EBS volume configurations, and spot/on-demand fleet composition changes. Tune Spark runtime configurations at the cluster level, including executor memory/core ratios, YARN container sizing, dynamic resource allocation settings, memory overhead parameters, and shuffle service configurations, to achieve optimal memory utilization without impacting job SLAs. Perform custom operations and iterative experiments using internal tooling to validate optimization impact: own end-to-end deployment, test execution, metric validation, and derive actionable insights from results. Collaborate with service teams to review cluster architectures, discuss findings, propose optimization plans, and align resolution strategies while communicating effectively across engineering leadership and technical stakeholders. Monitor service health metrics and troubleshoot operational issues during and after optimization activities, ensuring zero degradation to job completion times, data processing throughput, and downstream SLAs. Develop comprehensive operational runbooks, SOPs, documentation, and technical specifications that capture cluster optimization patterns and can be consumed by both human engineers and AI agents to orchestrate optimization workflows at scale. Extract scalable learnings from optimization engagements and develop programmatic frameworks that enable the initiative to scale across hundreds of EMR clusters, including training and enabling other vendor engineers to execute optimization playbooks. Job Description Must Have Technical/Functional Skills AWS EMR Cluster Operations Spark & YARN Tuning Memory Optimization & Capacity Analysis EC2 Right-Sizing & Cost Optimization Monitoring & Troubleshooting (CloudWatch/Logs) Roles & Responsibilities Analyze EMR cluster metrics, Spark application telemetry, and YARN resource utilization to identify over-provisioned memory allocations, underutilized executors, and suboptimal cluster configurations that contribute to excessive DRAM consumption. Recommend and implement cluster-level optimizations, including instance family right-sizing (e.g., migrating from memory-optimized R-type to compute-optimized C-type instances), node count adjustments, EBS volume configurations, and spot/on-demand fleet composition changes. Tune Spark runtime configurations at the cluster level, including executor memory/core ratios, YARN container sizing, dynamic resource allocation settings, memory overhead parameters, and shuffle service configurations, to achieve optimal memory utilization without impacting job SLAs. Perform custom operations and iterative experiments using internal tooling to validate optimization impact: own end-to-end deployment, test execution, metric validation, and derive actionable insights from results. Collaborate with service teams to review cluster architectures, discuss findings, propose optimization plans, and align resolution strategies while communicating effectively across engineering leadership and technical stakeholders. Monitor service health metrics and troubleshoot operational issues during and after optimization activities, ensuring zero degradation to job completion times, data processing throughput, and downstream SLAs. Develop comprehensive operational runbooks, SOPs, documentation, and technical specifications that capture cluster optimization patterns and can be consumed by both human engineers and AI agents to orchestrate optimization workflows at scale. Extract scalable learnings from optimization engagements and develop programmatic frameworks that enable the initiative to scale across hundreds of EMR clusters, including training and enabling other vendor engineers to execute optimization playbooks. TCS Employee Benefits Summary Discretionary Annual Incentive. Comprehensive Medical Coverage: Medical & Health, Dental & Vision, Disability Planning & Insurance, Pet Insurance Plans. Family Support: Maternal & Parental Leaves. Insurance Options: Aut& Home Insurance, Identity Theft Protection. Convenience & Professional Growth: Commuter Benefits & Certification & Training Reimbursement. Time Off: Vacation, Time Off, Sick Leave & Holidays. Legal & Financial Assistance: Legal Assistance, 401K Plan, Performance Bonus, College Fund, Student Loan Refinancing. Salary Range-$120,000-$160,000 a year #J-18808-Ljbffr

Before you apply

Applying takes about a minute. These four things decide how fast it moves after that.

Your profile is current

It's what we read first. Occupations, seniority and locations matter more than a long history.

Two examples you can talk through

Not a portfolio — just two pieces of work where you can explain the decisions and what you'd change.

A number in mind

What you're on now and what would make you move. We negotiate better when we know both.

Your notice period

Employers plan around it, and it's the question that stalls offers most often.

Once you apply, someone reads it and calls you before anything reaches the employer — usually within two working days.

More like this