AI Research Engineer, Pre-Training
objective paradigmChicago (IL)
AI Research Engineer, Pre-Training
Posted 17 days ago
objective paradigmChicago (IL)
SENIORITY
Senior
About the role
Reference: 26517
Location: Chicago, IL, United States
Industry: Trading Firm
Posted: 2026-08-17
Contact: Ethan Hudson
Email: ehudson@oprecruiting.comPhone: 13176505687
Job Title: AI Research Engineer – Large-Scale Pre-Training
Location: New York, NY or London, UKAbout the Opportunity
A leading quantitative financial technology enterprise is seeking an exceptional AI Research Engineer to advance its large-scale pre-training capabilities. In this role, you will help architect, optimize, and deploy massive foundation models designed to process extensive market and alternative data streams. If you want your engineering contributions to directly drive high-impact trading outcomes while leveraging extensive GPU infrastructure, this role offers an ideal environment with zero prior financial background required.
Responsibilities:
Optimize all phases of distributed deep learning pre-training, focusing on networking performance, memory management, data pipelines, and fault-tolerant execution.
Collaborate directly with machine learning researchers to co-design network architectures and establish long-term compute roadmaps.
Write high-performance lower-level code and custom kernels to maximize compute utilization across large-scale GPU infrastructure.
Adapt and translate cutting-edge deep learning methodologies from non-financial domains into production-grade trading models.
Requirements:
(must-have)2+ years of hands-on professional experience engineering deep learning systems across any technical field (e.g., robotics, physical sciences, computer vision, audio, or recommendation systems).Deep expertise in modern hardware acceleration and deep learning frameworks (such as PyTorch, JAX, CUDA, Triton, or specialized compiler DSLs).Proven track record of developing or tuning low-level training infrastructure, custom kernels, or distributed training workflows (e.g., parallelism techniques, CUDA Graphs, XLA).Ability to solve open-ended systems performance challenges without relying on off-the-shelf software packages. No prior experience in finance or quantitative trading is necessary. Preferred
Qualifications:
(nice-to-have) Practical experience building or fine-tuning Large Language Models (LLMs) at scale. Background in specialized hardware programming, FPGA/ASIC integration, or custom compute acceleration. Strong track record of cross-domain methodology transfer.
Compensation & Benefits:
Highly competitive base salary, generous performance-based bonuses, and comprehensive benefits. Access to exceptional GPU-per-engineer ratios and state-of-the-art compute resources.
Equal Opportunity Employer:
All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, or protected veteran status.#LI-EH1
Before you apply
Applying takes about a minute. These four things decide how fast it moves after that.
Your profile is current
It's what we read first. Occupations, seniority and locations matter more than a long history.
Two examples you can talk through
Not a portfolio — just two pieces of work where you can explain the decisions and what you'd change.
A number in mind
What you're on now and what would make you move. We negotiate better when we know both.
Your notice period
Employers plan around it, and it's the question that stalls offers most often.
Once you apply, someone reads it and calls you before anything reaches the employer — usually within two working days.
More like this
