Responsibilities : • Should have 7 years of experience with a strong foundation in ML inference, deployment, and quality validation. Should be capable of end-to-end ownership from model deployment ...
Responsibilities : • Should have 7 years of experience with a strong foundation in ML inference, deployment, and quality validation. Should be capable of end-to-end ownership from model deployment ...
Partner with ML, Product, Infrastructure, and Cloud teams to translate product and business ... Preferred : • Experience with ML inference infrastructure, model serving systems, or GPU ...
Partner with ML, Product, Infrastructure, and Cloud teams to translate product and business ... Preferred : • Experience with ML inference infrastructure, model serving systems, or GPU ...
$88K - $106K/yr
You will combine deep ML expertise, systems engineering, and performance analysis to deliver faster, more efficient AI experiences. * Optimize machine learning inference systems to improve latency ...
$88K - $106K/yr
You will combine deep ML expertise, systems engineering, and performance analysis to deliver faster, more efficient AI experiences. * Optimize machine learning inference systems to improve latency ...
Staff AI Inference and Acceleration Engineer
San Jose, CA · On-site
$180K - $275K/yr
Partner closely with the AI/ML team to define model architecture constraints that are hardware ... Deep understanding of AI/ML inference - model formats (ONNX, TFLite, etc.), inference runtimes, and ...
Staff AI Inference and Acceleration Engineer
San Jose, CA · On-site
$180K - $275K/yr
Partner closely with the AI/ML team to define model architecture constraints that are hardware ... Deep understanding of AI/ML inference - model formats (ONNX, TFLite, etc.), inference runtimes, and ...
Staff Embedded ML Engineer, Edge AI
$142K - $187K/yr
As a key contributor, you will lead the on-device inference and performance optimization of ML models powering outdoor monitoring in the home security space. This role is less about inventing new CV ...
Quick apply
Staff Embedded ML Engineer, Edge AI
$142K - $187K/yr
As a key contributor, you will lead the on-device inference and performance optimization of ML models powering outdoor monitoring in the home security space. This role is less about inventing new CV ...
AI / Embedded ML Engineer
Saratoga, CA · On-site
$145K - $190K/yr
... inference latency • Use frameworks including TensorFlow Lite Micro, Edge Impulse, ONNX Runtime, and ExecuTorch • Integrate ML inference into embedded firmware written in C, C++, or Rust • ...
AI / Embedded ML Engineer
Saratoga, CA · On-site
$145K - $190K/yr
... inference latency • Use frameworks including TensorFlow Lite Micro, Edge Impulse, ONNX Runtime, and ExecuTorch • Integrate ML inference into embedded firmware written in C, C++, or Rust • ...
AI / Embedded ML Engineer
Saratoga, CA · Hybrid
$150K - $225K/yr
Software Embedding and Systems Integration ◦ Write clean, well-tested embedded software that integrates ML inference into real-time systems ◦ Work with RTOS environments such as FreeRTOS and ...
Quick apply
AI / Embedded ML Engineer
Saratoga, CA · Hybrid
$150K - $225K/yr
Software Embedding and Systems Integration ◦ Write clean, well-tested embedded software that integrates ML inference into real-time systems ◦ Work with RTOS environments such as FreeRTOS and ...
Staff Embedded ML Engineer, Edge AI
$142K - $187K/yr
As a key contributor, you will lead the on-device inference and performance optimization of ML models powering outdoor monitoring in the home security space. This role is less about inventing new CV ...
Staff Embedded ML Engineer, Edge AI
$142K - $187K/yr
As a key contributor, you will lead the on-device inference and performance optimization of ML models powering outdoor monitoring in the home security space. This role is less about inventing new CV ...
AI / Embedded ML Engineer
Saratoga, CA · On-site
$145K - $190K/yr
... inference latency • Use frameworks including TensorFlow Lite Micro, Edge Impulse, ONNX Runtime, and ExecuTorch • Integrate ML inference into embedded firmware written in C, C++, or Rust • ...
AI / Embedded ML Engineer
Saratoga, CA · On-site
$145K - $190K/yr
... inference latency • Use frameworks including TensorFlow Lite Micro, Edge Impulse, ONNX Runtime, and ExecuTorch • Integrate ML inference into embedded firmware written in C, C++, or Rust • ...
Senior Software Engineer, ML Evaluation Infra and Efficiency
Mountain View, CA · On-site
$123K - $169K/yr
Improve runtime goodput of ML inference workload and efficiency of metrics computations, ensuring scalability and reliability across distributed environments. * Implement and maintain advanced ML ...
Senior Software Engineer, ML Evaluation Infra and Efficiency
Mountain View, CA · On-site
$123K - $169K/yr
Improve runtime goodput of ML inference workload and efficiency of metrics computations, ensuring scalability and reliability across distributed environments. * Implement and maintain advanced ML ...
Staff Embedded ML Engineer, Edge AI
Boston, MA · On-site
$142K - $187K/yr
As a key contributor, you will lead the on-device inference and performance optimization of ML models powering outdoor monitoring in the home security space. This role is less about inventing new CV ...
Staff Embedded ML Engineer, Edge AI
Boston, MA · On-site
$142K - $187K/yr
As a key contributor, you will lead the on-device inference and performance optimization of ML models powering outdoor monitoring in the home security space. This role is less about inventing new CV ...
About the team The AI/ML Engineering team builds and operates ClickHouse's AI and machine learning ... Design and implement AI-powered features across the full stack, from backend inference services to ...
About the team The AI/ML Engineering team builds and operates ClickHouse's AI and machine learning ... Design and implement AI-powered features across the full stack, from backend inference services to ...
Senior ML Ops Engineer
Columbus, OH · On-site
$100K - $138K/yr
Own the reliability and scalability of ML inference infrastructure. Design and tune autoscaling policies against real production traffic patterns, implement rate limiting and backpressure mechanisms ...
Senior ML Ops Engineer
Columbus, OH · On-site
$100K - $138K/yr
Own the reliability and scalability of ML inference infrastructure. Design and tune autoscaling policies against real production traffic patterns, implement rate limiting and backpressure mechanisms ...
Senior ML Ops Engineer
$100K - $138K/yr
Own the reliability and scalability of ML inference infrastructure. Design and tune autoscaling policies against real production traffic patterns, implement rate limiting and backpressure mechanisms ...
Senior ML Ops Engineer
$100K - $138K/yr
Own the reliability and scalability of ML inference infrastructure. Design and tune autoscaling policies against real production traffic patterns, implement rate limiting and backpressure mechanisms ...
Senior Software Engineer - Wilmington, Ma - ML/Ai/Robotics
Wilmington, MA · On-site
$133K - $176K/yr
... engineers to define inference contracts, latency budgets, interface requirements, and safe ... using ML inference runtimes such as TFLite, ONNX, or TensorRT within control or perception ...
New
Senior Software Engineer - Wilmington, Ma - ML/Ai/Robotics
Wilmington, MA · On-site
$133K - $176K/yr
... engineers to define inference contracts, latency budgets, interface requirements, and safe ... using ML inference runtimes such as TFLite, ONNX, or TensorRT within control or perception ...
New
Software Engineer - GenAI inference
San Francisco, CA · On-site
$142K - $204K/yr
Solid understanding of ML inference internals: attention, MLPs, recurrent modules, quantization, sparse operations, etc. * Hands-on experience with CUDA, GPU programming, and key libraries (cuBLAS ...
Software Engineer - GenAI inference
San Francisco, CA · On-site
$142K - $204K/yr
Solid understanding of ML inference internals: attention, MLPs, recurrent modules, quantization, sparse operations, etc. * Hands-on experience with CUDA, GPU programming, and key libraries (cuBLAS ...
Senior Software Engineer (Typescript / FrontEnd) - AI/ML
OR · Remote
$122K - $161K/yr
About the team The AI/ML Engineering team builds and operates ClickHouse's AI and machine learning ... Design and implement AI-powered features across the full stack, from backend inference services to ...
Senior Software Engineer (Typescript / FrontEnd) - AI/ML
OR · Remote
$122K - $161K/yr
About the team The AI/ML Engineering team builds and operates ClickHouse's AI and machine learning ... Design and implement AI-powered features across the full stack, from backend inference services to ...
Founding AI/ML Engineer
San Francisco, CA · On-site
$125K - $200K/yr
LLM/VLM inference; distributed backend infrastructure. Requirements * 4-6+ years of engineering experience, ideally 2+ in applied ML * Strong Python backend plus ML-ops familiarity (FastAPI ...
Quick apply
Founding AI/ML Engineer
San Francisco, CA · On-site
$125K - $200K/yr
LLM/VLM inference; distributed backend infrastructure. Requirements * 4-6+ years of engineering experience, ideally 2+ in applied ML * Strong Python backend plus ML-ops familiarity (FastAPI ...
GCP AI/ML Engineer
Chicago, IL · On-site
Design automated workflows for data ingestion, feature engineering, model training, evaluation, and inference. * Orchestrate ML workflows using Python, Vertex AI, BigQuery, and Cloud Storage.
GCP AI/ML Engineer
Chicago, IL · On-site
Design automated workflows for data ingestion, feature engineering, model training, evaluation, and inference. * Orchestrate ML workflows using Python, Vertex AI, BigQuery, and Cloud Storage.
Staff Software Engineer - GenAI inference
San Francisco, CA · On-site
$190K - $232K/yr
Deep understanding of ML inference internals: attention, MLPs, recurrent modules, quantization, sparse operations, etc. * Hands-on experience with CUDA, GPU programming, and key libraries (cuBLAS ...
Staff Software Engineer - GenAI inference
San Francisco, CA · On-site
$190K - $232K/yr
Deep understanding of ML inference internals: attention, MLPs, recurrent modules, quantization, sparse operations, etc. * Hands-on experience with CUDA, GPU programming, and key libraries (cuBLAS ...
Ml Inference information
See salary details
$37.5K - $52K
2% of jobs
$52K - $66.4K
3% of jobs
$66.4K - $80.9K
6% of jobs
$80.9K - $95.3K
9% of jobs
$100K is the 25th percentile. Wages below this are outliers.
$95.3K - $109.8K
15% of jobs
The median wage is $119.4K / yr.
$109.8K - $124.2K
22% of jobs
$132.2K is the 75th percentile. Wages above this are outliers.
$124.2K - $138.7K
32% of jobs
$138.7K - $153.1K
3% of jobs
$153.1K - $167.6K
4% of jobs
$167.6K - $182K
1% of jobs
$182K - $196.5K
2% of jobs
$37.5K
$122.7K
$196.5K
How much do ml inference jobs pay per year?
What is ML inference?
What is the difference between Ml Inference vs Data Scientist?
| Aspect | ML Inference | Data Scientist |
|---|---|---|
| Required Credentials | Knowledge of machine learning models, programming skills | Degree in data science, statistics, or related fields |
| Work Environment | Deploying models in production, real-time data processing | Data analysis, model development, research |
| Industry Usage | AI product deployment, software companies | Research institutions, tech firms, consulting |
ML Inference focuses on deploying trained models to make predictions on new data, often in real-time. Data Scientists develop and analyze models, working primarily in research and development. While both roles require understanding of machine learning, ML Inference emphasizes deployment and operationalization, whereas Data Scientists focus on model creation and analysis.
What are some common challenges faced by ML inference engineers when deploying models to production?
What are the key skills and qualifications needed to thrive in ML inference?
Is ML inference a high paying job?

Job description
eTeam is seeking a Machine Learning Platform Engineer to work onsite in Richardson, TX. The role requires strong expertise in ML inference, deployment, and quality validation, with responsibilities including end-to-end ownership from model deployment to user impact.
Responsibilities:
• Should have 7 years of experience with a strong foundation in ML inference, deployment, and quality validation. Should be capable of end-to-end ownership from model deployment to user impact, with the ability to quickly adapt to new technologies.
• Should have strong expertise in ML benchmarking and collaboration, along with hands-on experience deploying models on cloud platforms, preferably GCP. Familiarity with Java/JVM-based systems for model integration, streaming data architectures, and hybrid (on-prem and cloud) environments is essential.
• Must possess solid system design and distributed systems knowledge for troubleshooting, and hands-on experience with ML frameworks such as TensorFlow, PyTorch, or JAX.
Qualifications:
Required:
• 7 years of experience with a strong foundation in ML inference, deployment, and quality validation.
• Capability of end-to-end ownership from model deployment to user impact.
• Ability to quickly adapt to new technologies.
• Strong expertise in ML benchmarking and collaboration.
• Hands-on experience deploying models on cloud platforms, preferably GCP.
• Familiarity with Java/JVM-based systems for model integration.
• Experience with streaming data architectures.
• Experience in hybrid (on-prem and cloud) environments.
• Solid system design and distributed systems knowledge for troubleshooting.
• Hands-on experience with ML frameworks such as TensorFlow, PyTorch, or JAX.
Company:
eTeam is a staffing agency that also provides payrolling services. Founded in 1999, the company is headquartered in Somerset, USA, with a team of 501-1000 employees. The company is currently Late Stage.