1

Ml Inference Jobs in Alameda, CA (NOW HIRING)

ML Infrastructure Engineer

San Mateo, CA · On-site

$122K - $160K/yr

Own inference and model-serving infrastructure end to end, from design through production ... Collaborate closely with ML and infrastructure teams to ensure seamless integration and performance ...

next page

Showing results 1-20

Ml Inference information

See Alameda, CA salary details

$42.5K

$139.1K

$222.7K

How much do ml inference jobs pay per year?

As of Aug 29, 2026, the average yearly pay for ml inference in Alameda, CA is $139,107.00, according to ZipRecruiter salary data. Most workers in this role earn between $111,600.00 and $154,100.00 per year, depending on experience, location, and employer.

What is ML inference?

ML inference refers to the process of using a trained machine learning model to make predictions or decisions based on new data. After a model has been trained on historical data, inference is the phase where that model is deployed and used in real-world applications, such as recognizing speech, detecting objects in images, or recommending products. The focus in ML inference is on speed, efficiency, and scalability to ensure quick predictions, often in real time. This process is critical for practical applications like mobile apps, web services, and embedded systems. Optimizing inference involves reducing latency, memory usage, and computational requirements.

What are the key skills and qualifications needed to thrive in ML inference?

To thrive in ML Inference, you need a solid background in machine learning principles, programming (Python or C++), and experience with deploying models at scale, often supported by a degree in computer science or a related field. Familiarity with frameworks and tools such as TensorFlow, PyTorch, ONNX, and cloud platforms like AWS SageMaker or Google AI Platform is typically required. Strong problem-solving skills, attention to detail, and effective communication are crucial soft skills for collaborating with multidisciplinary teams and optimizing model performance. These skills ensure efficient, scalable, and reliable deployment of machine learning solutions in real-world applications.

What are some common challenges faced by ML inference engineers when deploying models to production?

ML Inference Engineers often encounter challenges such as optimizing model latency and throughput to meet production requirements, ensuring compatibility with diverse hardware environments, and managing model versioning and updates without disrupting service. Additionally, balancing resource utilization and inference accuracy while monitoring real-time performance metrics is crucial. Collaboration with data scientists, DevOps, and software engineers is typically essential to streamline deployment and maintain robust, scalable inference pipelines.

What is the difference between Ml Inference vs Data Scientist?

AspectML InferenceData Scientist
Required CredentialsKnowledge of machine learning models, programming skillsDegree in data science, statistics, or related fields
Work EnvironmentDeploying models in production, real-time data processingData analysis, model development, research
Industry UsageAI product deployment, software companiesResearch institutions, tech firms, consulting

ML Inference focuses on deploying trained models to make predictions on new data, often in real-time. Data Scientists develop and analyze models, working primarily in research and development. While both roles require understanding of machine learning, ML Inference emphasizes deployment and operationalization, whereas Data Scientists focus on model creation and analysis.

What job categories do people searching Ml Inference jobs in Alameda, CA look for?

The top searched job categories for Ml Inference jobs in Alameda, CA are:

What cities near Alameda, CA are hiring for Ml Inference jobs?

Cities near Alameda, CA with the most Ml Inference job openings:

Software Engineer, ML Inference Platform

Nuro

Mountain View, CA • On-site

$160K - $240K/yr

Full-time

Posted 17 days ago


Job description

Who We Are
Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that's why we're building a universal autonomy platform: self-driving for all roads and all rides.
Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles.
With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected.
Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors.
About the Role
The ML Infrastructure team is responsible for building & improving the core infrastructure for autonomy teams at Nuro. In this role, you will work closely with teams across Nuro, to design, build and deploy core infrastructure components in machine learning model life cycle, to push the autonomous future forward. You will have an opportunity to work across the full stack of machine learning solutions - from designing robust and scalable model & data pipelines to building to deploying the optimized models on Nuro's fleet of self-driving robots!
About the Work
  • Design and develop ML workflow pipelines to train, optimize, validate, and deploy Nuro autonomy models.
  • Develop and maintain continuous testing and monitoring systems for core ML infrastructure components.
  • Develop observability to track ML model lifecycles from data generation to on-road validation.
  • Maintain an in-house ML inference platform to serve large language models efficiently.
  • Maintain an in-house ML compiler platform to compile, deploy, and validate Nuro autonomy models to various hardware platforms.

About You
  • You have a degree in BS, MS.c or Ph.D with 1+ years of relevant experience
  • Strong proficiency in Python or similar languages. Familiarity with c++.
  • Technical excellence: Develop with a high standard for performance, scalability, and code quality.
  • Domain experience: Experience working with ML pipelines; ability to understand and design complex distributed systems
  • Independence and productivity: Ability and willingness to dive into unfamiliar territory, learn fast, and improve complex systems.

Bonus Points
  • Strong proficiency in C++ or other high-performance low-level languages
  • Experience working with large-scale distributed systems
  • Experience with system & framework design
  • Experience with data workflow orchestration platforms
  • Experience with ML compilers

At Nuro, your base pay is one part of your total compensation package. For this position, the reasonably expected base pay range is between $160,360 and $240,540 for the level at which this job has been scoped. Your base pay will depend on several factors, including your experience, qualifications, education, location, and skills. In the event that you are considered for a different level, a higher or lower pay range would apply. This position is also eligible for an annual performance bonus, equity, and a competitive benefits package.
At Nuro, we celebrate differences and are committed to a diverse workplace that fosters inclusion and psychological safety for all employees. Nuro is proud to be an equal opportunity employer and expressly prohibits any form of workplace discrimination based on race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other legally protected characteristics.