Machine Learning Research Engineer
$200,000 - $300,000 per year
ApplyMachine Learning Research Engineer
Posted yesterday
SENIORITY
Lead
SALARY
$200,000 - $300,000 per year
About the role
- Serve as the primary feedback loop for the entire ML stack.
- Actively run complex models through our full ML pipeline to comprehensively test both the training and inference environments.
- Validate the central infrastructure in practice, seeing exactly how new research ideas fare and identifying system bottlenecks before broader rollout to research teams.
- Streamline Rapid Prototyping for ML Research:
- Build high-level abstractions that allow users to bypass setup friction.
- Integrate our core ML tooling directly with our underlying simulation and data frameworks, providing a unified entry point to access our full tech stack.
- Enable rapid iteration on real-world data and seamless distributed training via Ray.
- Agentic Workflows for ML Research:
- Leverage AI agents and auto-research workflows to autonomously generate experiments, stress-test our distributed clusters, and provide data-driven, actionable feedback on what infrastructure needs to be optimized or built next.
- Research Platform Feedback & Insights Sharing:
- Act as the critical bridge between infrastructure builders and ML researchers.
- Be the first to exhaustively test new models and push the platform's limits.
- Document and publish empirical findings on system capabilities and hardware performance.
- Take your validated insights to assist engineering teams with platform improvements and advise researchers on how to best leverage the stack.
- Deep proficiency in Python and software design principles.
- Ability to build clean, scalable APIs and abstractions that other developers and researchers are enthusiastic about using.
- Applied Machine Learning:
- Hands-on experience with modern frameworks (PyTorch, TensorFlow, etc.)Strong practical understanding of how to train, evaluate, and deploy models at scale.
- Distributed Compute:
- Experience scaling ML workloads across GPUs and multi-node clusters using frameworks like Ray, Dask, or PyTorch Distributed.
- AI Agent Workflows:
- Familiarity with LLM tooling, agentic frameworks, and using AI to automate coding, research, or testing tasks.
- System Profiling & Optimization:
- Ability to debug and identify bottlenecks across hardware and software layers (e.g., memory limits, GPU utilization, data pipeline latency).Comp: $200-300K + Bonus
Before you apply
Applying takes about a minute. These four things decide how fast it moves after that.
Your profile is current
It's what we read first. Occupations, seniority and locations matter more than a long history.
Two examples you can talk through
Not a portfolio — just two pieces of work where you can explain the decisions and what you'd change.
A number in mind
What you're on now and what would make you move. We negotiate better when we know both.
Your notice period
Employers plan around it, and it's the question that stalls offers most often.
Once you apply, someone reads it and calls you before anything reaches the employer — usually within two working days.
More like this
