Engineer - Cluster Operations
Posted yesterday
tata consultancy servicesSeattle (WA)
Computer Systems Engineers/ArchitectsComputer Systems Design Services
SENIORITY
Senior
SALARY
$120,000-$160,000 a year
About the role
Monitoring & Troubleshooting (CloudWatch/Logs)
Analyze EMR cluster metrics, Spark application telemetry, and YARN resource utilization to identify over-provisioned memory allocations, underutilized executors, and suboptimal cluster configurations that contribute to excessive DRAM consumption.
Recommend and implement cluster-level optimizations, including instance family right-sizing (e.g., migrating from memory-optimized R-type to compute-optimized C-type instances), node count adjustments, EBS volume configurations, and spot/on-demand fleet composition changes.
Tune Spark runtime configurations at the cluster level, including executor memory/core ratios, YARN container sizing, dynamic resource allocation settings, memory overhead parameters, and shuffle service configurations, to achieve optimal memory utilization without impacting job SLAs.
Perform custom operations and iterative experiments using internal tooling to validate optimization impact: own end-to-end deployment, test execution, metric validation, and derive actionable insights from results.
Collaborate with service teams to review cluster architectures, discuss findings, propose optimization plans, and align resolution strategies while communicating effectively across engineering leadership and technical stakeholders.
Monitor service health metrics and troubleshoot operational issues during and after optimization activities, ensuring zero degradation to job completion times, data processing throughput, and downstream SLAs.
Develop comprehensive operational runbooks, SOPs, documentation, and technical specifications that capture cluster optimization patterns and can be consumed by both human engineers and AI agents to orchestrate optimization workflows at scale.
Extract scalable learnings from optimization engagements and develop programmatic frameworks that enable the initiative to scale across hundreds of EMR clusters, including training and enabling other vendor engineers to execute optimization playbooks.
Job Description
Must Have Technical/Functional Skills
AWS EMR Cluster Operations
Spark & YARN Tuning
Memory Optimization & Capacity Analysis
EC2 Right-Sizing & Cost Optimization
Monitoring & Troubleshooting (CloudWatch/Logs)
Roles & Responsibilities
Analyze EMR cluster metrics, Spark application telemetry, and YARN resource utilization to identify over-provisioned memory allocations, underutilized executors, and suboptimal cluster configurations that contribute to excessive DRAM consumption.
Recommend and implement cluster-level optimizations, including instance family right-sizing (e.g., migrating from memory-optimized R-type to compute-optimized C-type instances), node count adjustments, EBS volume configurations, and spot/on-demand fleet composition changes.
Tune Spark runtime configurations at the cluster level, including executor memory/core ratios, YARN container sizing, dynamic resource allocation settings, memory overhead parameters, and shuffle service configurations, to achieve optimal memory utilization without impacting job SLAs.
Perform custom operations and iterative experiments using internal tooling to validate optimization impact: own end-to-end deployment, test execution, metric validation, and derive actionable insights from results.
Collaborate with service teams to review cluster architectures, discuss findings, propose optimization plans, and align resolution strategies while communicating effectively across engineering leadership and technical stakeholders.
Monitor service health metrics and troubleshoot operational issues during and after optimization activities, ensuring zero degradation to job completion times, data processing throughput, and downstream SLAs.
Develop comprehensive operational runbooks, SOPs, documentation, and technical specifications that capture cluster optimization patterns and can be consumed by both human engineers and AI agents to orchestrate optimization workflows at scale.
Extract scalable learnings from optimization engagements and develop programmatic frameworks that enable the initiative to scale across hundreds of EMR clusters, including training and enabling other vendor engineers to execute optimization playbooks.
TCS Employee Benefits Summary
Discretionary Annual Incentive.
Comprehensive Medical Coverage: Medical & Health, Dental & Vision, Disability Planning & Insurance, Pet Insurance Plans.
Family Support: Maternal & Parental Leaves.
Insurance Options: Aut& Home Insurance, Identity Theft Protection.
Convenience & Professional Growth: Commuter Benefits & Certification & Training Reimbursement.
Time Off: Vacation, Time Off, Sick Leave & Holidays.
Legal & Financial Assistance: Legal Assistance, 401K Plan, Performance Bonus, College Fund, Student Loan Refinancing.
Salary Range-$120,000-$160,000 a year
#J-18808-Ljbffr
Before you apply
Applying takes about a minute. These four things decide how fast it moves after that.
Your profile is current
It's what we read first. Occupations, seniority and locations matter more than a long history.
Two examples you can talk through
Not a portfolio — just two pieces of work where you can explain the decisions and what you'd change.
A number in mind
What you're on now and what would make you move. We negotiate better when we know both.
Your notice period
Employers plan around it, and it's the question that stalls offers most often.
Once you apply, someone reads it and calls you before anything reaches the employer — usually within two working days.
More like this
