Machine Learning Engineer

Posted today

virtu financialBrooklyn (NY)
Computer Systems Engineers/ArchitectsComputing Infrastructure Providers, Data Processing, Web Hosting, and Related Services

SENIORITY

Senior

Apply

About the role

Overview In this ML infrastructure role, you’ll build the tools and systems that power quantitative researchers, enabling rapid experimentation and scalable ML across GPU clusters. You’ll shape how ML is used at Virtu by designing data and compute pipelines, experiment tracking, and reproducibility features. You’ll collaborate with quants and engineers to reduce friction in research workflows and improve efficiency. This is a hands-on, impact-driven position at the intersection of ML, trading, and systems engineering, with a clear path to shape the firm’s research infrastructure. Compensation / Benefitscompetitive salaryequal opportunity employer Responsibilities Design and build experiment tracking, job orchestration, and reproducibility infrastructure for rapid iteration and reliable comparisons Develop tools covering the simulation lifecycle, including back-tests and production monitoring, and enhance simulators Own visibility into GPU cluster utilization to detect bottlenecks and optimize compute usage Diagnose and resolve performance issues across training pipelines (data loading, I/O, GPU utilization, inter-node communication)Build and maintain data pipelines moving financial data into training workflows with strong correctness/versioning Develop feature storage/retrieval for fast, reproducible access to training data at scale Collaborate with researchers to reduce workflow friction via tooling and infrastructure improvements Coordinate with infrastructure engineers on capacity planning, cloud/on-prem tradeoffs, and tooling decisions Stay current with ML infrastructure developments and bring valuable tools into the stack Key requirements 5+ years of experience in ML engineering, research infrastructure, or HPC environments Strong Python engineering skills with maintainable, tested code Exposure to C++ in performance-sensitive contexts (a plus)Experience building/operating distributed training infrastructure and knowledge of collective communication libraries (NCCL, Horovod)Practical experience with experiment tracking systems and strong opinions on good research infrastructure Comfort with Linux systems stack (storage, networking, job scheduling)Excellent communication skills for cross-disciplinary collaboration Intellectually curious and self-driven with proactive problem solvingstrong collaborationclear communicationproblem solving mindset PythonC++ (performance-sensitive contexts)Distributed training infrastructure

Before you apply

Applying takes about a minute. These four things decide how fast it moves after that.

Your profile is current

It's what we read first. Occupations, seniority and locations matter more than a long history.

Two examples you can talk through

Not a portfolio — just two pieces of work where you can explain the decisions and what you'd change.

A number in mind

What you're on now and what would make you move. We negotiate better when we know both.

Your notice period

Employers plan around it, and it's the question that stalls offers most often.

Once you apply, someone reads it and calls you before anything reaches the employer — usually within two working days.

More like this