The role involves prototyping various algorithms suitable for inference hardware and guiding the hardware team on product definition. Responsibilities : • Prototype and optimize emerging ML ...
The role involves prototyping various algorithms suitable for inference hardware and guiding the hardware team on product definition. Responsibilities : • Prototype and optimize emerging ML ...
Partner with ML, Product, Infrastructure, and Cloud teams to translate product and business ... Preferred : • Experience with ML inference infrastructure, model serving systems, or GPU ...
Partner with ML, Product, Infrastructure, and Cloud teams to translate product and business ... Preferred : • Experience with ML inference infrastructure, model serving systems, or GPU ...
Software Engineer - GenAI inference
San Francisco, CA · On-site
$142.20 - $204.60/hr
Solid understanding of ML inference internals: attention, MLPs, recurrent modules, quantization, sparse operations, etc. * Hands-on experience with CUDA, GPU programming, and key libraries (cuBLAS ...
Software Engineer - GenAI inference
San Francisco, CA · On-site
$142.20 - $204.60/hr
Solid understanding of ML inference internals: attention, MLPs, recurrent modules, quantization, sparse operations, etc. * Hands-on experience with CUDA, GPU programming, and key libraries (cuBLAS ...
AI / Embedded ML Engineer
Saratoga, CA · On-site
$145K - $190K/yr
... inference latency • Use frameworks including TensorFlow Lite Micro, Edge Impulse, ONNX Runtime, and ExecuTorch • Integrate ML inference into embedded firmware written in C, C++, or Rust • ...
AI / Embedded ML Engineer
Saratoga, CA · On-site
$145K - $190K/yr
... inference latency • Use frameworks including TensorFlow Lite Micro, Edge Impulse, ONNX Runtime, and ExecuTorch • Integrate ML inference into embedded firmware written in C, C++, or Rust • ...
AI / Embedded ML Engineer
Saratoga, CA · On-site
$145K - $190K/yr
Medical
PTO
... inference latency • Use frameworks including TensorFlow Lite Micro, Edge Impulse, ONNX Runtime, and ExecuTorch • Integrate ML inference into embedded firmware written in C, C++, or Rust • ...
AI / Embedded ML Engineer
Saratoga, CA · On-site
$145K - $190K/yr
Medical
PTO
... inference latency • Use frameworks including TensorFlow Lite Micro, Edge Impulse, ONNX Runtime, and ExecuTorch • Integrate ML inference into embedded firmware written in C, C++, or Rust • ...
AI / Embedded ML Engineer
Saratoga, CA · Hybrid
$150K - $225K/yr
Software Embedding and Systems Integration ◦ Write clean, well-tested embedded software that integrates ML inference into real-time systems ◦ Work with RTOS environments such as FreeRTOS and ...
Quick apply
AI / Embedded ML Engineer
Saratoga, CA · Hybrid
$150K - $225K/yr
Software Embedding and Systems Integration ◦ Write clean, well-tested embedded software that integrates ML inference into real-time systems ◦ Work with RTOS environments such as FreeRTOS and ...
Drive LLM/VLM inference optimization, including batching, scheduling, quantization, model serving ... Mentor engineers in distributed systems, scalable ML infrastructure, and multimodal AI engineering ...
Drive LLM/VLM inference optimization, including batching, scheduling, quantization, model serving ... Mentor engineers in distributed systems, scalable ML infrastructure, and multimodal AI engineering ...
Senior Software Engineer, ML Evaluation Infra and Efficiency
Mountain View, CA · On-site
$145K - $192K/yr
... inference speed and resource utilization. • Collaborate with ML engineers to understand evaluation requirements and scenarios, and improve DevX of the evaluation infrastructure. • Improve runtime ...
Senior Software Engineer, ML Evaluation Infra and Efficiency
Mountain View, CA · On-site
$145K - $192K/yr
... inference speed and resource utilization. • Collaborate with ML engineers to understand evaluation requirements and scenarios, and improve DevX of the evaluation infrastructure. • Improve runtime ...
Senior Software Engineer, ML Evaluation Infra and Efficiency
Mountain View, CA · On-site
$123K - $169K/yr
Improve runtime goodput of ML inference workload and efficiency of metrics computations, ensuring scalability and reliability across distributed environments. * Implement and maintain advanced ML ...
Senior Software Engineer, ML Evaluation Infra and Efficiency
Mountain View, CA · On-site
$123K - $169K/yr
Improve runtime goodput of ML inference workload and efficiency of metrics computations, ensuring scalability and reliability across distributed environments. * Implement and maintain advanced ML ...
Software Engineer - GenAI inference
San Francisco, CA · On-site
$142K - $204K/yr
Solid understanding of ML inference internals: attention, MLPs, recurrent modules, quantization, sparse operations, etc. * Hands-on experience with CUDA, GPU programming, and key libraries (cuBLAS ...
Software Engineer - GenAI inference
San Francisco, CA · On-site
$142K - $204K/yr
Solid understanding of ML inference internals: attention, MLPs, recurrent modules, quantization, sparse operations, etc. * Hands-on experience with CUDA, GPU programming, and key libraries (cuBLAS ...
Senior Software Engineer, ML Platform
San Francisco, CA · Remote
$144K - $190K/yr
Expand batch ML inference systems across scheduling, parallelism, cost controls, observability, failure handling, and rollback * Own and improve the feature store, including offline and online ...
New
Quick apply
Senior Software Engineer, ML Platform
San Francisco, CA · Remote
$144K - $190K/yr
Expand batch ML inference systems across scheduling, parallelism, cost controls, observability, failure handling, and rollback * Own and improve the feature store, including offline and online ...
New
Senior Software Engineer, ML Platform
San Francisco, CA · Remote
$144K - $190K/yr
Expand batch ML inference systems across scheduling, parallelism, cost controls, observability, failure handling, and rollback * Own and improve the feature store, including offline and online ...
New
Senior Software Engineer, ML Platform
San Francisco, CA · Remote
$144K - $190K/yr
Expand batch ML inference systems across scheduling, parallelism, cost controls, observability, failure handling, and rollback * Own and improve the feature store, including offline and online ...
New
Senior Software Engineer, ML Platform
San Francisco, CA · On-site
$144K - $190K/yr
Expand batch ML inference systems across scheduling, parallelism, cost controls, observability, failure handling, and rollback * Own and improve the feature store, including offline and online ...
Senior Software Engineer, ML Platform
San Francisco, CA · On-site
$144K - $190K/yr
Expand batch ML inference systems across scheduling, parallelism, cost controls, observability, failure handling, and rollback * Own and improve the feature store, including offline and online ...
Founding AI/ML Engineer
San Francisco, CA · On-site
$125K - $200K/yr
LLM/VLM inference; distributed backend infrastructure. Requirements * 4-6+ years of engineering experience, ideally 2+ in applied ML * Strong Python backend plus ML-ops familiarity (FastAPI ...
Quick apply
Founding AI/ML Engineer
San Francisco, CA · On-site
$125K - $200K/yr
LLM/VLM inference; distributed backend infrastructure. Requirements * 4-6+ years of engineering experience, ideally 2+ in applied ML * Strong Python backend plus ML-ops familiarity (FastAPI ...
Staff Software Engineer - GenAI inference
San Francisco, CA · On-site
$190K - $232K/yr
Deep understanding of ML inference internals: attention, MLPs, recurrent modules, quantization, sparse operations, etc. * Hands-on experience with CUDA, GPU programming, and key libraries (cuBLAS ...
Staff Software Engineer - GenAI inference
San Francisco, CA · On-site
$190K - $232K/yr
Deep understanding of ML inference internals: attention, MLPs, recurrent modules, quantization, sparse operations, etc. * Hands-on experience with CUDA, GPU programming, and key libraries (cuBLAS ...
We are looking for engineers with ML software & systems expertise to help build the next generation Waymo onboard ML inference engine for Waymo fundamental model. You'll work across the entire ML ...
We are looking for engineers with ML software & systems expertise to help build the next generation Waymo onboard ML inference engine for Waymo fundamental model. You'll work across the entire ML ...
Robotics Autonomy ML Engineer - VLA & Manipulation
San Francisco, CA · On-site
$150K - $200K/yr
Medical
Dental
Vision
Build reliable, high‑speed robot autonomy software stack optimized for inference performance ... Background in real‑time ML inference systems, simulation‑to‑reality transfer, or advanced ...
Robotics Autonomy ML Engineer - VLA & Manipulation
San Francisco, CA · On-site
$150K - $200K/yr
Medical
Dental
Vision
Build reliable, high‑speed robot autonomy software stack optimized for inference performance ... Background in real‑time ML inference systems, simulation‑to‑reality transfer, or advanced ...
Staff Software Engineer - GenAI inference
San Francisco, CA · On-site
$190.90 - $232.80/hr
Deep understanding of ML inference internals: attention, MLPs, recurrent modules, quantization, sparse operations, etc. * Hands‑on experience with CUDA, GPU programming, and key libraries (cuBLAS ...
Staff Software Engineer - GenAI inference
San Francisco, CA · On-site
$190.90 - $232.80/hr
Deep understanding of ML inference internals: attention, MLPs, recurrent modules, quantization, sparse operations, etc. * Hands‑on experience with CUDA, GPU programming, and key libraries (cuBLAS ...
Member of Technical Staff, ML Systems
Santa Clara, CA · On-site
Medical
Dental
Vision
Life
Retirement
Prototype and optimize emerging ML inference systems. * Develop novel memory models for expandable vRAM. * Write efficient GPU kernels for data movement. * Perform design-space exploration ...
Member of Technical Staff, ML Systems
Santa Clara, CA · On-site
Medical
Dental
Vision
Life
Retirement
Prototype and optimize emerging ML inference systems. * Develop novel memory models for expandable vRAM. * Write efficient GPU kernels for data movement. * Perform design-space exploration ...
Required : • 3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems. • Deep experience with GPU programming and performance work ...
Required : • 3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems. • Deep experience with GPU programming and performance work ...
Ml Inference information
What is ML inference?
What is the difference between Ml Inference vs Data Scientist?
| Aspect | ML Inference | Data Scientist |
|---|---|---|
| Required Credentials | Knowledge of machine learning models, programming skills | Degree in data science, statistics, or related fields |
| Work Environment | Deploying models in production, real-time data processing | Data analysis, model development, research |
| Industry Usage | AI product deployment, software companies | Research institutions, tech firms, consulting |
ML Inference focuses on deploying trained models to make predictions on new data, often in real-time. Data Scientists develop and analyze models, working primarily in research and development. While both roles require understanding of machine learning, ML Inference emphasizes deployment and operationalization, whereas Data Scientists focus on model creation and analysis.
What are some common challenges faced by ML inference engineers when deploying models to production?
What are the key skills and qualifications needed to thrive in ML inference?
Is ML inference a high paying job?
What are popular job titles related to Ml Inference jobs in California?
For Ml Inference jobs in California, the most frequently searched job titles are:
What job categories do people searching Ml Inference jobs in California look for?
The top searched job categories for Ml Inference jobs in California are:
What cities in California are hiring for Ml Inference jobs?
Cities in California with the most Ml Inference job openings:

Full-time
Re-posted 24 days ago
Job description
Netpreme is seeking a motivated LLM Systems Engineer to explore new and unconventional inference systems based on emerging hardware. The role involves prototyping various algorithms suitable for inference hardware and guiding the hardware team on product definition.
Responsibilities:
• Prototype and optimize emerging ML inference systems.
• Develop novel memory models for expandable vRAM.
• Write efficient GPU kernels for data movement.
• Perform design-space exploration, implementation, and benchmarking of inference engines, both in simulations and on real hardware.
Qualifications:
Required:
• MS or PhD in computer systems, ideally with a focus on LLM inference and/or distributed systems.
• Prior experience contributing to the core LLM inference infrastructures (vLLM, SGLang, TensorRT, etc.).
• Prior experience in accelerator programming (e.g. CUDA, JAX/Pallas, ROCm).
Preferred:
• Advanced computer architectures and performance engineering skills is a big plus.
Company:
Netpreme develops networked memory tiering technology that expands accelerator memory capacity for AI workloads. Founded in 2024, the company is headquartered in Cambridge, USA, with a team of 11-50 employees. The company is currently Early Stage.