Software Engineer 3, Platform

Posted yesterday

digitasNeedham Heights (MA)

SENIORITY

Lead

Apply

About the role

Software Engineer 3 Welcome to Our World We've been leading the charge in the affiliate industry from day one—establishing performance marketing and paving the way for future innovations. We're known for maintaining one of the largest, most reliable partnership platforms with impeccable, personalized service. Founded in Santa Barbara, California in 1998, CJ (formerly Commission Junction) stands as the most trusted name in performance marketing. We specialize in building partnerships between top brands and reputable publishers to drive revenue and business growth. CJ's industry-leading solutions make us the platform of choice for over 3,800 global brands across sectors like retail, travel, finance, technology, and home services. As part of Publicis Groupe, our savvy data capabilities, cutting-edge tech, and strategic expertise facilitate genuine connections, allowing brands to reach consumers wherever they are. A Quick Peek at Affiliate Marketing Think back to your last online purchase. Did an influencer tip you off about a great product and offer a discount? Or perhaps you relied on a trusted review site to make your decision? Whatever path you took, affiliate publishers likely played a role by influencing, informing, or helping you find the best deal. CJ connects brands with these publishers, creating valuable resources for shoppers like you.
Overview: You must be work authorized in the United States without the need for employer sponsorship. This is a hybrid role requiring 3 days a week in office. About CJ Engineering Are we a good match for you? At CJ, we are passionate about software engineering. We build exciting software, with quality and maintainability in mind. We believe in common sense, simplicity, and efficiency. We practice critical thinking, challenge each other no matter the title, and believe in the wisdom of the team. That makes us Engineers and not just developers. Here are some of the principles that set CJ engineers apart. Engineering autonomy: Here, business decisions are made by business people, and technical decisions are made by technical people. Full stack: Expect to be involved and to gain competence in every aspect of software engineering, from frontend, to database, to requirements analysis, to testing, to helping choose technologies. Clean, maintainable code: Code is read more often than it is written, so we put in the effort to write it well in the first place. Pairing: The highest quality code is produced by close collaboration, so we pair by default. TDD: Quality is baked into our process through Test Driven Development Ownership: Engineers own the full lifecycle of what they build—from design and implementation to deployment, monitoring, production support, and on-call rotations. If we build it, we support it. Operational Excellence: We embrace Infrastructure as Code, CI/CD, automation, and observability to build reliable systems and deliver software safely, efficiently, and at scale. We believe in Agile values, and incremental development. We constantly experiment, retrospect, and adjust. We are committed to finding out how AI can amplify our productivity. We view AI as a force multiplier, not a replacement for good engineering. If this sounds exciting, we want to hear from you! As a Software Engineer 3 on the Engineering Experience (Eng Exp) platform team, you help run and evolve the platform that powers CJ's production systems across multiple AWS regions. "Platform" here is broad - it is the Kubernetes clusters, but also the observability stack every squad depends on, the CI/CD and artifact infrastructure their builds run through, the AWS networking that connects them, the secrets and access systems that gate them, and the cost visibility that keeps them accountable. Eng Exp owns all of it. This is not just an infrastructure role - your value is in engineering judgment: how you evaluate systems, detect risk, and make decisions under uncertainty. You'll own meaningful pieces of these systems independently and drive changes from design through production. We want real depth in the systems below, not just familiarity with the tool names.
Responsibilities: What systems you will work on: Observability & monitoring — Prometheus, Alertmanager, Grafana, and Open Telemetry across production regions. This is not dashboard-building: you'll own cardinality budgets and recording-rule design, keep a production Prometheus healthy as it outgrows a single shard (federation / sharding / long-term store), and understand Alertmanager HA and the blast radius of alert-routing config. Deep Prometheus and Alertmanager knowledge is a core requirement. Kubernetes & cloud infrastructure — multi-region EKS clusters: upgrades, node group and Karpenter management, controller lifecycle, and add-on / configuration management. Spot failure modes before they happen (subnet IP exhaustion, API server latency, ArgoCD reconciliation lag, Prometheus cardinality, Karpenter consolidation disruption).AWS networking — VPC and subnet design, CIDR management, VPC peering, Route 53, security groups, and NAT gateway topology across accounts and regions, plus the 24/7 networking alarms for prod networking between clusters and squad resources. CI/CD & artifact management — Git Lab administration (runner fleet, cache, access - not just pipeline authoring), Git Ops delivery through ArgoCD, and the Nexus artifact repository including its storage lifecycle as it grows. Access & identity — Vault secrets management, IAM roles and service accounts for apps in clusters, cluster permission management for audit compliance, and AI model access management. Turn recurring access requests into self-service workflows that are hard to misuse. Cost observability — Open Cost, EBS orphan cleanup, cost anomaly investigation, and rightsizing attribution across teams.
What You'll Do: Own and evolve meaningful pieces of the systems above - with a focus on what is happening and why Manage infrastructure-as-code with Terraform across AWS accounts Build and maintain Git Lab CI/CD pipelines and Git Ops delivery (ArgoCD)Help enforce platform standards: RBAC, admission webhooks, resource limits, Limit Ranges Turn recurring requests (ingress, DNS, service accounts) into self-service workflows Respond to and help drive resolution of platform incidents, focused on learning and system improvement Act as a reviewer of infrastructure changes - Terraform, Kubernetes configs, observability config Technologies We Use Kubernetes / EKS (multi-cluster, multi-region), Karpenter, cert-manager, external-dns Prometheus, Alertmanager, Grafana, Open Telemetry (and long-term storage / sharding for Prometheus) AWS networking (VPC, VPC peering, Transit Gateway, Route 53, NAT Gateway, security groups, subnet/CIDR design across accounts and regions) Terraform, AWS (IAM, EKS, S3, EBS) ArgoCD, Git Lab CI/CD, Nexus (artifact registry), Docker, container image build pipelines Vault, Open Cost Kubernetes controllers/operators (reconciliation patterns, restart safety) - Go experience is a plus, not required
Qualifications
What We Look For: 3+ years of experience in software or infrastructure engineering Bachelor's degree or equivalent experience Hands-on production experience with Kubernetes and at least one major cloud (AWS preferred)Real operational depth in at least one system we own beyond the cluster - most valuably the observability stack (Prometheus/Alertmanager at scale), but AWS networking, Vault, or artifact/CI infrastructure also count. We are filtering for people who have run these systems, not just used them. Comfortable owning infrastructure-as-code (Terraform) and CI/CD pipelines Can reason about tradeoffs and communicate the pros and cons of multiple approaches Effective communication; thrives in a collaborative, pair-friendly team culture
Nice to Have: AWS networking depth (Transit Gateway, multi-account topology)Prometheus long-term storage / sharding (Thanos, Cortex, Mimir, or equivalent)Kubernetes controllers/operators - Go experience is a plus, not required Policy-as-code (Kyverno / OPA)
What Success Looks Like: You own platform components independently and ship changes that are intentional and low-risk The systems you own are understood deeply enough that we stop making decisions we have to reverse Production issues are understood quickly because of the observability and instincts you bring Engineers can deploy and debug services with less platform intervention over time

Before you apply

Applying takes about a minute. These four things decide how fast it moves after that.

Your profile is current

It's what we read first. Occupations, seniority and locations matter more than a long history.

Two examples you can talk through

Not a portfolio — just two pieces of work where you can explain the decisions and what you'd change.

A number in mind

What you're on now and what would make you move. We negotiate better when we know both.

Your notice period

Employers plan around it, and it's the question that stalls offers most often.

Once you apply, someone reads it and calls you before anything reaches the employer — usually within two working days.

More like this