Deep understanding of ML inference internals: attention, MLPs, recurrent modules, quantization, sparse operations, etc. * Hands-on experience with CUDA, GPU programming, and key libraries (cuBLAS ...
Deep understanding of ML inference internals: attention, MLPs, recurrent modules, quantization, sparse operations, etc. * Hands-on experience with CUDA, GPU programming, and key libraries (cuBLAS ...
... of ML inference internals: attention, MLPs, recurrent modules, quantization, sparse operations, etc. • Hands-on experience with CUDA, GPU programming, and key libraries (cuBLAS, cuDNN, NCCL, etc ...
... of ML inference internals: attention, MLPs, recurrent modules, quantization, sparse operations, etc. • Hands-on experience with CUDA, GPU programming, and key libraries (cuBLAS, cuDNN, NCCL, etc ...
AI/ML Platform Engineer
Alexandria, VA · On-site
FastAPI and microservices for ML inference * InfrastructureasCode (Terraform) * Kubernetes and Docker for scalable ML workloads * Distributed/cloud systems design with AWS * Edgetocloud system ...
AI/ML Platform Engineer
Alexandria, VA · On-site
FastAPI and microservices for ML inference * InfrastructureasCode (Terraform) * Kubernetes and Docker for scalable ML workloads * Distributed/cloud systems design with AWS * Edgetocloud system ...
Apple's Server ML Frameworks team in GPU, Graphics and Machine Learning works on enabling Apple Intelligence through high-performance, distributed inference of GenAI applications (such as LLMs) on ...
Apple's Server ML Frameworks team in GPU, Graphics and Machine Learning works on enabling Apple Intelligence through high-performance, distributed inference of GenAI applications (such as LLMs) on ...
Senior Software Development Engineer, AI/ML, AWS Neuron, Model Inference
Cupertino, CA · On-site
$128K - $177K/yr
The Inference Enablement and Acceleration team is at the forefront of running a wide range of models and supporting novel architecture alongside maximizing their performance for AWS's custom ML ...
Senior Software Development Engineer, AI/ML, AWS Neuron, Model Inference
Cupertino, CA · On-site
$128K - $177K/yr
The Inference Enablement and Acceleration team is at the forefront of running a wide range of models and supporting novel architecture alongside maximizing their performance for AWS's custom ML ...
ML Software Engineer
$142K - $263K/yr
Our team builds ML-inference applications and services on Apple Silicon in the datacenter, specifically focusing in recent years on generative AI as part of the Private Cloud Compute component of ...
ML Software Engineer
$142K - $263K/yr
Our team builds ML-inference applications and services on Apple Silicon in the datacenter, specifically focusing in recent years on generative AI as part of the Private Cloud Compute component of ...
Member of Technical Staff - ML Systems & Inference
San Francisco, CA · On-site
$150K - $350K/yr
About the role Gimlet is seeking a Member of Technical Staff focused on ML Systems and Inference. In this role, you will design and build the inference systems that execute full models end-to-end ...
Member of Technical Staff - ML Systems & Inference
San Francisco, CA · On-site
$150K - $350K/yr
About the role Gimlet is seeking a Member of Technical Staff focused on ML Systems and Inference. In this role, you will design and build the inference systems that execute full models end-to-end ...
Senior Software Development Engineer, AI/ML, AWS Neuron, Model Inference
Cupertino, CA · On-site
$128K - $177K/yr
The Inference Enablement and Acceleration team is at the forefront of running a wide range of models and supporting novel architecture alongside maximizing their performance for AWS's custom ML ...
Senior Software Development Engineer, AI/ML, AWS Neuron, Model Inference
Cupertino, CA · On-site
$128K - $177K/yr
The Inference Enablement and Acceleration team is at the forefront of running a wide range of models and supporting novel architecture alongside maximizing their performance for AWS's custom ML ...
Senior Software Development Engineer, AI/ML, AWS Neuron, Model Inference
Cupertino, CA · On-site
$128K - $177K/yr
The Inference Enablement and Acceleration team is at the forefront of running a wide range of models and supporting novel architecture alongside maximizing their performance for AWS's custom ML ...
Senior Software Development Engineer, AI/ML, AWS Neuron, Model Inference
Cupertino, CA · On-site
$128K - $177K/yr
The Inference Enablement and Acceleration team is at the forefront of running a wide range of models and supporting novel architecture alongside maximizing their performance for AWS's custom ML ...
Senior Software Engineer, ML Platform
San Francisco, CA · On-site
$220K - $265K/yr
Decompose data scientist training/inference notebooks into reusable, tested components (libraries, pipelines, templates) with clear interfaces and documentation. * Create developer-friendly ML ...
Senior Software Engineer, ML Platform
San Francisco, CA · On-site
$220K - $265K/yr
Decompose data scientist training/inference notebooks into reusable, tested components (libraries, pipelines, templates) with clear interfaces and documentation. * Create developer-friendly ML ...
$89K - $123K/yr
About the Role We are seeking an experienced Senior ML Inference Engineer to join our team, focusing on optimizing and deploying our production virtual staining models at scale. The ideal candidate ...
$89K - $123K/yr
About the Role We are seeking an experienced Senior ML Inference Engineer to join our team, focusing on optimizing and deploying our production virtual staining models at scale. The ideal candidate ...
AI / Embedded ML Engineer
Saratoga, CA · On-site
$145K - $190K/yr
... inference latency • Use frameworks including TensorFlow Lite Micro, Edge Impulse, ONNX Runtime, and ExecuTorch • Integrate ML inference into embedded firmware written in C, C++, or Rust • ...
AI / Embedded ML Engineer
Saratoga, CA · On-site
$145K - $190K/yr
... inference latency • Use frameworks including TensorFlow Lite Micro, Edge Impulse, ONNX Runtime, and ExecuTorch • Integrate ML inference into embedded firmware written in C, C++, or Rust • ...
Senior Software Engineer, ML Platform
San Francisco, CA · Remote
$220K - $265K/yr
Decompose data scientist training/inference notebooks into reusable, tested components (libraries, pipelines, templates) with clear interfaces and documentation. * Create developer-friendly ML ...
Quick apply
Senior Software Engineer, ML Platform
San Francisco, CA · Remote
$220K - $265K/yr
Decompose data scientist training/inference notebooks into reusable, tested components (libraries, pipelines, templates) with clear interfaces and documentation. * Create developer-friendly ML ...
ML Framework (MetalLM) Engineer
$150K - $225K/yr
Work on cutting-edge ML inference framework project and optimize code for efficient and scalable ML inference using distributed compute strategies such as data, tensor, pipeline and expert ...
ML Framework (MetalLM) Engineer
$150K - $225K/yr
Work on cutting-edge ML inference framework project and optimize code for efficient and scalable ML inference using distributed compute strategies such as data, tensor, pipeline and expert ...
AI / Embedded ML Engineer
Saratoga, CA · On-site
$145K - $190K/yr
... inference latency • Use frameworks including TensorFlow Lite Micro, Edge Impulse, ONNX Runtime, and ExecuTorch • Integrate ML inference into embedded firmware written in C, C++, or Rust • ...
AI / Embedded ML Engineer
Saratoga, CA · On-site
$145K - $190K/yr
... inference latency • Use frameworks including TensorFlow Lite Micro, Edge Impulse, ONNX Runtime, and ExecuTorch • Integrate ML inference into embedded firmware written in C, C++, or Rust • ...
AI / Embedded ML Engineer
Saratoga, CA · On-site
$145K - $190K/yr
... inference latency • Use frameworks including TensorFlow Lite Micro, Edge Impulse, ONNX Runtime, and ExecuTorch • Integrate ML inference into embedded firmware written in C, C++, or Rust • ...
AI / Embedded ML Engineer
Saratoga, CA · On-site
$145K - $190K/yr
... inference latency • Use frameworks including TensorFlow Lite Micro, Edge Impulse, ONNX Runtime, and ExecuTorch • Integrate ML inference into embedded firmware written in C, C++, or Rust • ...
Staff AI Inference and Acceleration Engineer
$180K - $275K/yr
Partner closely with the AI/ML team to define model architecture constraints that are hardware ... Deep understanding of AI/ML inference - model formats (ONNX, TFLite, etc.), inference runtimes, and ...
Staff AI Inference and Acceleration Engineer
$180K - $275K/yr
Partner closely with the AI/ML team to define model architecture constraints that are hardware ... Deep understanding of AI/ML inference - model formats (ONNX, TFLite, etc.), inference runtimes, and ...
AI / Embedded ML Engineer
Saratoga, CA · Hybrid
$145K - $190K/yr
... inference latency Use frameworks including TensorFlow Lite Micro, Edge Impulse, ONNX Runtime, and ExecuTorch Integrate ML inference into embedded firmware written in C, C++, or Rust Profile and ...
AI / Embedded ML Engineer
Saratoga, CA · Hybrid
$145K - $190K/yr
... inference latency Use frameworks including TensorFlow Lite Micro, Edge Impulse, ONNX Runtime, and ExecuTorch Integrate ML inference into embedded firmware written in C, C++, or Rust Profile and ...
The role involves prototyping various algorithms suitable for inference hardware and guiding the hardware team on product definition. Responsibilities : • Prototype and optimize emerging ML ...
The role involves prototyping various algorithms suitable for inference hardware and guiding the hardware team on product definition. Responsibilities : • Prototype and optimize emerging ML ...
AI/ML Engineer (MLops)
Woodbridge, NJ · On-site
Deploying real-time ML inference pipelines processing millions of records at high throughput * Experience building end to end automated MLOps capabilities along with model and feature drift ...
AI/ML Engineer (MLops)
Woodbridge, NJ · On-site
Deploying real-time ML inference pipelines processing millions of records at high throughput * Experience building end to end automated MLOps capabilities along with model and feature drift ...
Ml Inference information
See salary details
$37.5K - $52K
2% of jobs
$52K - $66.4K
3% of jobs
$66.4K - $80.9K
6% of jobs
$80.9K - $95.3K
9% of jobs
$100K is the 25th percentile. Wages below this are outliers.
$95.3K - $109.8K
15% of jobs
The median wage is $119.4K / yr.
$109.8K - $124.2K
22% of jobs
$132.2K is the 75th percentile. Wages above this are outliers.
$124.2K - $138.7K
32% of jobs
$138.7K - $153.1K
3% of jobs
$153.1K - $167.6K
4% of jobs
$167.6K - $182K
1% of jobs
$182K - $196.5K
2% of jobs
$37.5K
$122.7K
$196.5K
How much do ml inference jobs pay per year?
What is ML inference?
What is the difference between Ml Inference vs Data Scientist?
| Aspect | ML Inference | Data Scientist |
|---|---|---|
| Required Credentials | Knowledge of machine learning models, programming skills | Degree in data science, statistics, or related fields |
| Work Environment | Deploying models in production, real-time data processing | Data analysis, model development, research |
| Industry Usage | AI product deployment, software companies | Research institutions, tech firms, consulting |
ML Inference focuses on deploying trained models to make predictions on new data, often in real-time. Data Scientists develop and analyze models, working primarily in research and development. While both roles require understanding of machine learning, ML Inference emphasizes deployment and operationalization, whereas Data Scientists focus on model creation and analysis.
What are some common challenges faced by ML inference engineers when deploying models to production?
What are the key skills and qualifications needed to thrive in ML inference?
Is ML inference a high paying job?

Job description
P-1285
About This RoleAs a staff software engineer for GenAI inference, you will lead the architecture, development, and optimization of the inference engine that powers Databricks Foundation Model API.. You'll bridge research advances and production demands, ensuring high throughput, low latency, and robust scaling. Your work will encompass the full GenAI inference stack: kernels, runtimes, orchestration, memory, and integration with frameworks and orchestration systems.
What You Will Do- Own and drive the architecture, design, and implementation of the inference engine, and collaborate on model-serving stack optimized for large-scale LLMs inference
- Partner closely with researchers to bring new model architectures or features (sparsity, activation compression, mixture-of-experts) into the engine
- Lead the end-to-end optimization for latency, throughput, memory efficiency, and hardware utilization across GPUs, and accelerators
- Define and guide standards to build and maintain instrumentation, profiling, and tracing tooling to uncover bottlenecks and guide optimizations
- Architect scalable routing, batching, scheduling, memory management, and dynamic loading mechanisms for inference workloads
- Ensure reliability, reproducibility, and fault tolerance in the inference pipelines, including A/B launches, rollback, and model versioning
- Collaborate cross-functionally on Integrating with federated, distributed inference infrastructure - orchestrate across nodes, balance load, handle communication overhead
- Drive cross-team collaboration: with platform engineers, cloud infrastructure, and security/compliance teams
- Represent the team externally through benchmarks, whitepapers, and open-source contributions
- BS/MS/PhD in Computer Science, or a related field
- Strong software engineering background (6+ years or equivalent) in performance-critical systems
- Proven track record of owning complex system components and driving architectural decisions end-to-end
- Deep understanding of ML inference internals: attention, MLPs, recurrent modules, quantization, sparse operations, etc.
- Hands-on experience with CUDA, GPU programming, and key libraries (cuBLAS, cuDNN, NCCL, etc.)
- Strong background in distributed systems design, including RPC frameworks, queuing, RPC batching, sharding, memory partitioning
- Demonstrated ability to uncover and solve performance bottlenecks across layers (kernel, memory, networking, scheduler)
- Experience building instrumentation, tracing, and profiling tools for ML models
- Ability to lead through influence - work closely with ML researchers, translate novel model ideas into production systems
- Excellent communication and leadership skills, with a proactive and ownership-driven mindset
- Bonus: published research or open-source contributions in ML systems, inference optimization, or model serving
About Databricks
Sourced by ZipRecruiter
Industry
Software development
Company size
5,001 - 10,000 Employees
Headquarters location
San Francisco, CA, US
Year founded
2013