Senior ML Engineer
Posted yesterday
cloudbolt softwareSilver Spring (MD)
Data ScientistsCustom Computer Programming Services
SENIORITY
Senior
About the role
Overview
You will own and advance the recommendation engine at the core of Storm Forge, Cloud Bolt’s Kubernetes resource optimization product. In a hands-on, production-focused ML role, you design time-series models, ensure data quality and safety, and collaborate cross-functionally to scale accurate, reliable recommendations. The work blends applied machine learning with production engineering to help customers optimize cloud spend without compromising workloads. You’ll shape the technical direction and raise the bar on performance and safety.
Compensation / Benefits Medical/Dental/Vision coverage 401k with Company Match Health & Dependent Care FSAUnlimited PTO11 Company Holidays Tuition Reimbursement`,`Paid Parental Leave
Responsibilities Own the end-to-end recommendation engine: model selection, algorithm design, preprocessing, and safe guardrails for live production workloads Design, evaluate, and productionize time-series forecasting and statistical models for right-sizing Kubernetes workloads across CPU, memory, GPU, and JVM heap Build and maintain data-quality layer to detect anomalies and filter telemetry before modeling Define and improve metrics and validation for recommendation quality (regression tests, behavioral validation, production metrics)Investigate and resolve customer-reported recommendation quality issues across data, preprocessing, and model behavior Serve as ML authority: guide technical direction, tradeoffs between models and heuristics, communicate to engineers and leadership Write production-grade Python for models and pipelines; share responsibility for surrounding service components (queues, caches, observability)Prototype and validate new optimization capabilities (new resource types, algorithms) from research to feature-flagged rollout Stay current on time-series forecasting and resource optimization techniques and evaluate practical applicability
Key requirements Master's degree or higher in a quantitative field 5+ years software engineering, with 3+ years operating ML or statistical systems in production Expert-level Python with strong typing and testing, fluency in numpy Hands-on experience with time-series analysis and forecasting; traditional statistical methods and anomaly detection Rigorous ML system testing: regression testing against baselines, behavioral validation, numerical reproducibility Working knowledge of Kubernetes: resource requests/limits, autoscaling, understanding failures under-provisioning Experience owning production services (queues, caches, observability, debugging from logs/metrics)Clear written and verbal communication to explain model behavior and tradeoffsstrong communicationcuriosity and rigorproblem-solving with a collaborative mindsettime-series forecasting Prophet or similar libraries Python (production-grade)
Before you apply
Applying takes about a minute. These four things decide how fast it moves after that.
Your profile is current
It's what we read first. Occupations, seniority and locations matter more than a long history.
Two examples you can talk through
Not a portfolio — just two pieces of work where you can explain the decisions and what you'd change.
A number in mind
What you're on now and what would make you move. We negotiate better when we know both.
Your notice period
Employers plan around it, and it's the question that stalls offers most often.
Once you apply, someone reads it and calls you before anything reaches the employer — usually within two working days.
More like this
