LLM Inference Engineer: Scale & Optimize Production
comunidade metodistaPalo Alto (CA)
LLM Inference Engineer: Scale & Optimize Production
Posted 2 days ago
comunidade metodistaPalo Alto (CA)
Computer Systems Engineers/ArchitectsComputing Infrastructure Providers, Data Processing, Web Hosting, and Related Services
SENIORITY
Lead
About the role
Hippocratic AI is seeking an experienced LLM Inference Engineer to optimize its LLM serving infrastructure in Palo Alto, CA. You will design multi-node architectures, implement multi-LoRA serving, and apply quantization techniques to reduce footprint while maintaining quality.
The role emphasizes practical production deployment, benchmarking, and latency-focused optimizations across GPUs, with a strong emphasis on Python, C++, and CUDA.
Before you apply
Applying takes about a minute. These four things decide how fast it moves after that.
Your profile is current
It's what we read first. Occupations, seniority and locations matter more than a long history.
Two examples you can talk through
Not a portfolio — just two pieces of work where you can explain the decisions and what you'd change.
A number in mind
What you're on now and what would make you move. We negotiate better when we know both.
Your notice period
Employers plan around it, and it's the question that stalls offers most often.
Once you apply, someone reads it and calls you before anything reaches the employer — usually within two working days.
More like this
