The Distributed LLM Inference Engineer will optimize systems for large-scale ML inference, working closely with product teams and the open-source community to deliver high-performance solutions.
The Distributed LLM Inference Engineer will optimize systems for large-scale ML inference, working closely with product teams and the open-source community to deliver high-performance solutions.
Software Engineer, ML Inference Platform
Mountain View, CA · On-site
$160K - $240K/yr
Maintain an in-house ML inference platform to serve large language models efficiently. * Maintain an in-house ML compiler platform to compile, deploy, and validate Nuro autonomy models to various ...
Software Engineer, ML Inference Platform
Mountain View, CA · On-site
$160K - $240K/yr
Maintain an in-house ML inference platform to serve large language models efficiently. * Maintain an in-house ML compiler platform to compile, deploy, and validate Nuro autonomy models to various ...
Gimlet Labs is building innovative solutions for AI workloads, and they are seeking a Member of Technical Staff focused on ML systems and inference. In this role, you will design and build inference ...
Gimlet Labs is building innovative solutions for AI workloads, and they are seeking a Member of Technical Staff focused on ML systems and inference. In this role, you will design and build inference ...
Software Engineer, ML Inference Platform
Mountain View, CA · On-site
$160K - $240K/yr
Maintain an in-house ML inference platform to serve large language models efficiently. * Maintain an in-house ML compiler platform to compile, deploy, and validate Nuro autonomy models to various ...
Quick apply
Software Engineer, ML Inference Platform
Mountain View, CA · On-site
$160K - $240K/yr
Maintain an in-house ML inference platform to serve large language models efficiently. * Maintain an in-house ML compiler platform to compile, deploy, and validate Nuro autonomy models to various ...
Distributed LLM Inference Engineer
San Francisco, CA · On-site
$170K - $245K/yr
Retirement
PTO
Familiarity with running ML inference at large scale with high throughput and low latency ... Familiarity with deep learning and deep learning frameworks (e.g. PyTorch) * Solid understanding of ...
Distributed LLM Inference Engineer
San Francisco, CA · On-site
$170K - $245K/yr
Retirement
PTO
Familiarity with running ML inference at large scale with high throughput and low latency ... Familiarity with deep learning and deep learning frameworks (e.g. PyTorch) * Solid understanding of ...
Software Engineer, ML Inference Platform
$160K - $240K/yr
Maintain an in-house ML inference platform to serve large language models efficiently. * Maintain an in-house ML compiler platform to compile, deploy, and validate Nuro autonomy models to various ...
Software Engineer, ML Inference Platform
$160K - $240K/yr
Maintain an in-house ML inference platform to serve large language models efficiently. * Maintain an in-house ML compiler platform to compile, deploy, and validate Nuro autonomy models to various ...
Staff Backend Engineer, ML Inference Systems
$244K - $317K/yr
Medical
Life
Retirement
PTO
Partner with ML engineers to ensure online serving infrastructure scales with growing model complexity and inference volumes, without compromising latency or throughput * Ensure the reliability ...
Staff Backend Engineer, ML Inference Systems
$244K - $317K/yr
Medical
Life
Retirement
PTO
Partner with ML engineers to ensure online serving infrastructure scales with growing model complexity and inference volumes, without compromising latency or throughput * Ensure the reliability ...
Staff Backend Engineer, ML Inference Systems
Mountain View, CA · On-site
$244K - $317K/yr
Medical
Life
Retirement
PTO
Partner with ML engineers to ensure online serving infrastructure scales with growing model complexity and inference volumes, without compromising latency or throughput * Ensure the reliability ...
Staff Backend Engineer, ML Inference Systems
Mountain View, CA · On-site
$244K - $317K/yr
Medical
Life
Retirement
PTO
Partner with ML engineers to ensure online serving infrastructure scales with growing model complexity and inference volumes, without compromising latency or throughput * Ensure the reliability ...
Senior Product Manager -AI/ML Inference Software
$149K - $197K/yr
As inference becomes the defining workload for production AI, you will help shape how AMD's software ecosystem enables efficient, reliable, and competitive large-scale model deployment. You will work ...
Senior Product Manager -AI/ML Inference Software
$149K - $197K/yr
As inference becomes the defining workload for production AI, you will help shape how AMD's software ecosystem enables efficient, reliable, and competitive large-scale model deployment. You will work ...
Work on cutting-edge ML inference framework project and optimize code for efficient and scalable ML inference using distributed compute strategies such as data, tensor, pipeline and expert ...
Work on cutting-edge ML inference framework project and optimize code for efficient and scalable ML inference using distributed compute strategies such as data, tensor, pipeline and expert ...
Member of Technical Staff, Inference
San Francisco, CA · On-site
Medical
Dental
Vision
San Francisco Description We're looking for an ML Inference Engineer with deep expertise in high-performance ML engineering. This is a highly technical, high-impact role focused on squeezing every ...
Member of Technical Staff, Inference
San Francisco, CA · On-site
Medical
Dental
Vision
San Francisco Description We're looking for an ML Inference Engineer with deep expertise in high-performance ML engineering. This is a highly technical, high-impact role focused on squeezing every ...
Sr. Software Engineer, ML Platform
San Francisco, CA · On-site
$144K - $190K/yr
Decompose data scientist training/inference notebooks into reusable, tested components (libraries, pipelines, templates) with clear interfaces and documentation. * Create developer-friendly ML ...
New
Sr. Software Engineer, ML Platform
San Francisco, CA · On-site
$144K - $190K/yr
Decompose data scientist training/inference notebooks into reusable, tested components (libraries, pipelines, templates) with clear interfaces and documentation. * Create developer-friendly ML ...
New
Solid understanding of ML inference internals: attention, MLPs, recurrent modules, quantization, sparse operations, etc. * Hands-on experience with CUDA, GPU programming, and key libraries (cuBLAS ...
Solid understanding of ML inference internals: attention, MLPs, recurrent modules, quantization, sparse operations, etc. * Hands-on experience with CUDA, GPU programming, and key libraries (cuBLAS ...
Work on cutting-edge ML inference framework project and optimize code for efficient and scalable ML inference using distributed compute strategies such as data, tensor, pipeline and expert ...
Work on cutting-edge ML inference framework project and optimize code for efficient and scalable ML inference using distributed compute strategies such as data, tensor, pipeline and expert ...
Senior Product Manager - ROCm & AI/ML Inference Software
Santa Clara, CA · On-site
$130 - $160/hr
... inference requirements and translates market signals into actionable product strategy. Open‑Source Community Engagement * Serve as AMD's active presence in the open‑source AI/ML community ...
Senior Product Manager - ROCm & AI/ML Inference Software
Santa Clara, CA · On-site
$130 - $160/hr
... inference requirements and translates market signals into actionable product strategy. Open‑Source Community Engagement * Serve as AMD's active presence in the open‑source AI/ML community ...
Senior Product Manager -AI/ML Inference Software
Santa Clara, CA · On-site
$179K/yr
... inference requirements and translates market signals into actionable product strategy. Open-Source Community Engagement * Serve as AMD's active presence in the open-source AI/ML community: monitor ...
Senior Product Manager -AI/ML Inference Software
Santa Clara, CA · On-site
$179K/yr
... inference requirements and translates market signals into actionable product strategy. Open-Source Community Engagement * Serve as AMD's active presence in the open-source AI/ML community: monitor ...
... ML inference internals: attention, MLPs, recurrent modules, quantization, sparse operations, etc ... • Hands-on experience with CUDA, GPU programming, and key libraries (cuBLAS, cuDNN, NCCL, etc ...
... ML inference internals: attention, MLPs, recurrent modules, quantization, sparse operations, etc ... • Hands-on experience with CUDA, GPU programming, and key libraries (cuBLAS, cuDNN, NCCL, etc ...
Maintain an in-house ML inference platform to serve large language models efficiently. * Maintain an in-house ML compiler platform to compile, deploy, and validate Nuro autonomy models to various ...
Maintain an in-house ML inference platform to serve large language models efficiently. * Maintain an in-house ML compiler platform to compile, deploy, and validate Nuro autonomy models to various ...
Deep understanding of ML inference internals: attention, MLPs, recurrent modules, quantization, sparse operations, etc. * Hands-on experience with CUDA, GPU programming, and key libraries (cuBLAS ...
Deep understanding of ML inference internals: attention, MLPs, recurrent modules, quantization, sparse operations, etc. * Hands-on experience with CUDA, GPU programming, and key libraries (cuBLAS ...
The Inference Enablement and Acceleration team is at the forefront of running a wide range of models and supporting novel architecture alongside maximizing their performance for AWS's custom ML ...
The Inference Enablement and Acceleration team is at the forefront of running a wide range of models and supporting novel architecture alongside maximizing their performance for AWS's custom ML ...
Ml Inference information
What is ML inference?
What is the difference between Ml Inference vs Data Scientist?
| Aspect | ML Inference | Data Scientist |
|---|---|---|
| Required Credentials | Knowledge of machine learning models, programming skills | Degree in data science, statistics, or related fields |
| Work Environment | Deploying models in production, real-time data processing | Data analysis, model development, research |
| Industry Usage | AI product deployment, software companies | Research institutions, tech firms, consulting |
ML Inference focuses on deploying trained models to make predictions on new data, often in real-time. Data Scientists develop and analyze models, working primarily in research and development. While both roles require understanding of machine learning, ML Inference emphasizes deployment and operationalization, whereas Data Scientists focus on model creation and analysis.
What are some common challenges faced by ML inference engineers when deploying models to production?
What are the key skills and qualifications needed to thrive in ML inference?
Is ML inference a high paying job?
What are popular job titles related to Ml Inference jobs in California?
For Ml Inference jobs in California, the most frequently searched job titles are:
What job categories do people searching Ml Inference jobs in California look for?
The top searched job categories for Ml Inference jobs in California are:
What cities in California are hiring for Ml Inference jobs?
Cities in California with the most Ml Inference job openings:

Job description
Anyscale is on a mission to democratize distributed computing and make it accessible to software developers. The Distributed LLM Inference Engineer will optimize systems for large-scale ML inference, working closely with product teams and the open-source community to deliver high-performance solutions.
Responsibilities:
• Iterate very quickly with product teams to ship the end to end solutions for Batch and Online inference at high scale which will be used by open-source Ray users and customers of Anyscale
• Work across the stack integrating Ray Data and LLM engine providing optimizations achieving low cost solutions for large scale ML inference
• Integrate with Open source software like vLLM, work closely with the community to adopt these techniques in Anyscale solutions, and also contribute improvements to open source
• Follow the latest state-of-the-art in the open source and the research community, implementing and extending best practices
Qualifications:
Required:
• Familiarity with running ML inference at large scale with high throughput and low latency
• Familiarity with deep learning and deep learning frameworks (e.g. PyTorch)
• Solid understanding of distributed systems, ML inference challenges
Preferred:
• ML Systems knowledge
• Experience using Ray
• Work closely with community on LLM engines like vLLM, TensorRT-LLM
• Contributions to deep learning frameworks (PyTorch, TensorFlow)
• Contributions to deep learning compilers (Triton, TVM, MLIR)
• Prior experience working on GPUs / CUDA
Company:
Anyscale develops a distributed computing platform that enables organizations to manage distributed workloads at scale. Founded in 2019, the company is headquartered in San Francisco, USA, with a team of 201-500 employees. The company is currently Growth Stage.
About Anyscale
Sourced by ZipRecruiter
Industry
Software development
Company size
51 - 200 Employees
Headquarters location
San Francisco, CA, US
Year founded
2019