1

Ml Inference Jobs in Phoenix, AZ (NOW HIRING)

Lead ML Ops Engineer

Tempe, AZ

$98K - $129K/yr

Oversee enterprisescale AI platforms supporting model training, inference, evaluation, monitoring ... Leadershiplevel expertise in AI/ML platform engineering, spanning MLOps, LLMOps, and AIOps.

Build and enhance AI/ML and GenAI-powered solutions using Python, LLMs, RAG, prompt engineering ... Flask, FastAPI, Hugging Face, Triton Inference Server * pgVector, Milvus, Prometheus, Elasticsearch ...

Explore and evaluate new AI/ML techniques, tools, and methodologies, applying relevant innovations ... and inference efficiency to minimize cost and latency while preserving accuracy. * MLOps ...

Lead AI Engineer

Phoenix, AZ · On-site

$99K - $131K/yr

LLM infrastructure, inference, and model gateways * Evaluation, observability, and safety tooling ... LangGraph, LangChain, AirFlow, etc Agentic AI and ML * Integration of commercial and open-source ...

Explore and evaluate new AI/ML techniques, tools, and methodologies, applying relevant innovations ... and inference efficiency to minimize cost and latency while preserving accuracy. * MLOps ...

Google AI Lead Architect

Tempe, AZ

$53 - $72.50/hr

Integrate and fine-tune Large Language Models (LLMs) and other AI/ML models into enterprise applications. Develop and implement strategies for model deployment, inference, and monitoring, with an ...

AI Solution Architect

Tempe, AZ · On-site

$60.25 - $79.50/hr

This individual will operate at the intersection of architecture, AI platform engineering, ML ... Real-time inference pipelines * Ensure architectural alignment with: * Cloud strategy * Enterprise ...

Experience with machine learning systems, including model training, evaluation, inference, and use of frameworks such as PyTorch, TensorFlow, or Scikit-learn. * Experience deploying AI/ML services ...

Experience with machine learning systems, including model training, evaluation, inference, and use of frameworks such asPyTorch, TensorFlow, or Scikit-learn. * Experience deploying AI/ML services ...

Integrate AI/ML models, LLMs, and Generative AI APIs into existing Java applications. Pipeline Engineering: Design and maintain data pipelines for machine learning model training and inference.

Experience serving ML models at scale using Triton Inference Server, TorchServe, Ray Serve, or similar. * Contributions to open-source ML projects or published research (papers, patents, or technical ...

Experience serving ML models at scale using Triton Inference Server, TorchServe, Ray Serve, or similar. * Contributions to open-source ML projects or published research (papers, patents, or technical ...

... inference techniques (uplift modeling, difference-in-differences, synthetic controls, instrumental variables) where randomized experiments aren't feasible * Translate business problems into ML ...

next page

Showing results 1-20

Ml Inference information

See Phoenix, AZ salary details

$37.2K

$121.9K

$195.1K

How much do ml inference jobs pay per year?

As of Jul 27, 2026, the average yearly pay for ml inference in Phoenix, AZ is $121,868.00, according to ZipRecruiter salary data. Most workers in this role earn between $97,800.00 and $135,000.00 per year, depending on experience, location, and employer.

What is a $900000 AI job?

A $900,000 AI job typically refers to high-level roles in artificial intelligence, such as senior machine learning engineers or AI research directors, often involving advanced skills in deep learning, data modeling, and programming with tools like Python and TensorFlow. These positions usually require extensive experience, specialized knowledge, and may include leadership responsibilities or strategic decision-making.

What is ML inference?

ML inference refers to the process of using a trained machine learning model to make predictions or decisions based on new data. After a model has been trained on historical data, inference is the phase where that model is deployed and used in real-world applications, such as recognizing speech, detecting objects in images, or recommending products. The focus in ML inference is on speed, efficiency, and scalability to ensure quick predictions, often in real time. This process is critical for practical applications like mobile apps, web services, and embedded systems. Optimizing inference involves reducing latency, memory usage, and computational requirements.

What is the difference between Ml Inference vs Data Scientist?

AspectML InferenceData Scientist
Required CredentialsKnowledge of machine learning models, programming skillsDegree in data science, statistics, or related fields
Work EnvironmentDeploying models in production, real-time data processingData analysis, model development, research
Industry UsageAI product deployment, software companiesResearch institutions, tech firms, consulting

ML Inference focuses on deploying trained models to make predictions on new data, often in real-time. Data Scientists develop and analyze models, working primarily in research and development. While both roles require understanding of machine learning, ML Inference emphasizes deployment and operationalization, whereas Data Scientists focus on model creation and analysis.

What engineer makes $500,000 a year?

Senior machine learning engineers with extensive experience, advanced skills in deep learning, and expertise in deploying large-scale models can earn salaries approaching or exceeding $500,000 annually, especially in high-cost-of-living areas or top tech companies. Compensation often includes base salary, bonuses, and stock options, reflecting their specialized knowledge and impact on product development.

Which 3 jobs will survive AI?

Jobs involving Ml Inference, such as data scientists, machine learning engineers, and AI system architects, are likely to persist as they require specialized expertise in developing, deploying, and maintaining AI models. These roles demand critical thinking, domain knowledge, and skills in programming and data analysis that are less easily automated. Continuous learning and staying updated with AI tools and frameworks are essential for these professions to remain relevant.

What are some common challenges faced by ML Inference Engineers when deploying models to production?

ML Inference Engineers often encounter challenges such as optimizing model latency and throughput to meet production requirements, ensuring compatibility with diverse hardware environments, and managing model versioning and updates without disrupting service. Additionally, balancing resource utilization and inference accuracy while monitoring real-time performance metrics is crucial. Collaboration with data scientists, DevOps, and software engineers is typically essential to streamline deployment and maintain robust, scalable inference pipelines.

Will MLE be replaced by AI?

Machine Learning Engineers (MLEs) design, develop, and optimize AI models and systems. While AI automation tools can assist with certain tasks, MLEs are essential for building, tuning, and maintaining complex models, making complete replacement unlikely in the near term. Their expertise in data handling, model deployment, and system integration remains critical in AI development environments.

What are the key skills and qualifications needed to thrive in ML Inference, and why are they important?

To thrive in ML Inference, you need a solid background in machine learning principles, programming (Python or C++), and experience with deploying models at scale, often supported by a degree in computer science or a related field. Familiarity with frameworks and tools such as TensorFlow, PyTorch, ONNX, and cloud platforms like AWS SageMaker or Google AI Platform is typically required. Strong problem-solving skills, attention to detail, and effective communication are crucial soft skills for collaborating with multidisciplinary teams and optimizing model performance. These skills ensure efficient, scalable, and reliable deployment of machine learning solutions in real-world applications.
What are popular job titles related to Ml Inference jobs in Phoenix, AZ? For Ml Inference jobs in Phoenix, AZ, the most frequently searched job titles are:
What job categories do people searching Ml Inference jobs in Phoenix, AZ look for? The top searched job categories for Ml Inference jobs in Phoenix, AZ are:
What cities near Phoenix, AZ are hiring for Ml Inference jobs? Cities near Phoenix, AZ with the most Ml Inference job openings:
Infographic showing various Ml Inference job openings in Phoenix, AZ as of July 2026, with employment types broken down into 100% Full Time. Highlights an 80% In-person, and 20% Remote job distribution, with an average salary of $121,868 per year, or $58.6 per hour.
IAM Engineer - Phoenix, Az

IAM Engineer - Phoenix, Az

Motion Recruitment

Phoenix, AZ • On-site

Other

Medical, Dental, Vision, Retirement, PTO

Posted 3 days ago


Job description


A leading financial services firm in Chandler, AZ is seeking a Senior Machine Learning Engineer to join their technology innovation team. In this position, you’ll be at the forefront of building state-of-the-art machine learning solutions, taking ownership of projects from the design phase through to on-device deployment.
You’ll architect and support robust pipelines for sensor data analysis, shaping ML systems that deliver real-time inference and anomaly detection on resource-constrained hardware. Collaborating closely with hardware, firmware, and platform teams, you’ll be responsible for seamlessly integrating and validating intelligent features within embedded environments. This role will involve building and automating MLOps solutions, overseeing experiment management, and ensuring scalable and compliant machine learning lifecycle practices suited to regulatory demands.
Key Responsibilities
  • Lead the implementation of sensor data pipelines—from requirements through real-world integration
  • Design, train, optimize, and deploy ML models specifically for edge or embedded platforms, focusing on performance and reliability
  • Combine deep experience in Python with one or more frameworks (PyTorch, TensorFlow) to deliver production-ready models
  • Partner with engineering and product teams to embed, monitor, and validate ML inference in device applications
  • Develop infrastructure and automations for experiment tracking, continuous training, evaluation, and secure model deployment
  • Maintain full documentation for all aspects of the ML lifecycle to enable audit, regulatory compliance, and operational excellence
Required Skills & Experience
  • Extensive experience working with sensor and streaming data in production settings
  • Expertise in Python and at least one major ML library (PyTorch, TensorFlow, or alternatives)
  • Practical experience deploying models on edge systems (e.g., TensorRT, ONNX, TFLite, or similar technologies)
  • Solid engineering background with proficiency in C or C++ for embedded systems collaboration
  • Familiarity with MLOps and model versioning, preferably in a regulated environment
Preferred Qualifications
  • 5+ years hands-on ML development, with significant experience in real-time or embedded applications
  • Background working on medical, wearable, or robotics platforms highly desirable
  • Demonstrated success working in cross-functional teams with hardware/firmware integration
  • Experience in high-growth companies or small teams
  • Bachelor’s or higher in Computer Science, Engineering, or a related field
Daily Activities
  • This role is hands-on and project-focused, where you’ll design new ML workflows, transform raw data, and continually iterate on model performance
  • You’ll support firmware engineers by ensuring seamless embedding and runtime validation of ML features
  • Work with DevOps and MLOps teams to automate deployments and create scalable, reliable ML solutions
  • Troubleshoot model issues, and refine algorithms for speed, accuracy, and compliance
Compensation & Benefits
Eligible for performance-based bonus or commission
Comprehensive benefits package including:
  • Health, dental, and vision insurance
  • Paid holidays and vacation
  • 401(k) retirement plan with company match (if applicable)
Applicants must have current authorization to work in the US on a full-time basis, now and in the future