Multimodal AI Model Optimization Research Engineer
tavusNew York (NY)
Multimodal AI Model Optimization Research Engineer
Posted 6 days ago
tavusNew York (NY)
Computer and Information Research ScientistsResearch and Development in the Physical, Engineering, and Life Sciences (except Nanotechnology and Biotechnology)
SENIORITY
Senior
About the role
We’re looking for an experienced Research Scientist/Engineer with a focus on model optimization to join our core AI team
Take cutting-edge research models and make them fast, efficient, and production-ready using sparsification, distillation, and quantization
Own the optimization lifecycle for key models: define metrics, run experiments, and benchmark trade-offs across latency, cost, and quality
Partner closely with researchers and engineers to turn new ideas into deployable systems
Benefits:
Comprehensive medical, dental and vision coverage for employees and their families, with 100% of monthly premiums paid by Tavus
Unlimited paid time off, as well as parental and pawental leave
Flexible hours
Remote-friendly, with teams at home or in co-working spaces across the globe
Debate and feedback are cornerstones of our culture
Yearly stipend for you to spend on any learning materials (i.e. newsletters, online courses, etc.)
Company retreats
Generous gear stipend
Frequent in person meet ups
Team members from diverse backgrounds – we’re looking for culture creators, not culture fits
Progressive, open-minded meritocracy
Qualifications:
Our ideal partner-in-crime thrives in startup environments, is comfortable prioritizing independently, and is willing to take calculated risks
We’re moving fast and looking for people who can help pave the path
Ability to read ML papers, reproduce results, and adapt ideas
Clear communication and collaboration skills
Strong understanding of inference performance and GPU/accelerator fundamentals
Strong Python coding skills and reliable research engineering practices
Experience working with large models and datasets in cloud environments
Hands-on experience with model optimization and compression, including knowledge distillation, pruning/sparsification, quantization, and mixed precision
Understanding of efficient architectures such as low-rank adapters
Strong experience in deep learning using PyTorch
Optimization of diffusion models, video/audio generative models, or large language models
Experience with real-time or streaming systems (low-latency APIs, WebRTC, streaming TTS/video)
Familiarity with TensorRT, ONNX Runtime, TVM, Triton, or XLA
Experience writing custom Triton/CUDA kernels or low-level performance tuning
Experience with experiment tracking, benchmarking, and profiling at scale
Prior experience in research engineering or applied science roles
Before you apply
Applying takes about a minute. These four things decide how fast it moves after that.
Your profile is current
It's what we read first. Occupations, seniority and locations matter more than a long history.
Two examples you can talk through
Not a portfolio — just two pieces of work where you can explain the decisions and what you'd change.
A number in mind
What you're on now and what would make you move. We negotiate better when we know both.
Your notice period
Employers plan around it, and it's the question that stalls offers most often.
Once you apply, someone reads it and calls you before anything reaches the employer — usually within two working days.
More like this
