1

Ml Inference Jobs in Arizona (NOW HIRING)

Deliver governed datasets and feature engineering/serving for ML training and real-time inference (online/offline consistency, caching, latency SLOs, backfills). A successful candidate would possess ...

Build and enhance AI/ML and GenAI-powered solutions using Python, LLMs, RAG, prompt engineering ... Flask, FastAPI, Hugging Face, Triton Inference Server * pgVector, Milvus, Prometheus, Elasticsearch ...

Lead AI Engineer

Phoenix, AZ · On-site

$99K - $131K/yr

LLM infrastructure, inference, and model gateways * Evaluation, observability, and safety tooling ... LangGraph, LangChain, AirFlow, etc Agentic AI and ML * Integration of commercial and open-source ...

Explore and evaluate new AI/ML techniques, tools, and methodologies, applying relevant innovations ... and inference efficiency to minimize cost and latency while preserving accuracy. * MLOps ...

Explore and evaluate new AI/ML techniques, tools, and methodologies, applying relevant innovations ... and inference efficiency to minimize cost and latency while preserving accuracy. * MLOps ...

Integrate and fine-tune Large Language Models (LLMs) and other AI/ML models into enterprise applications. Develop and implement strategies for model deployment, inference, and monitoring, with an ...

Google AI Lead Architect

Tempe, AZ

$53 - $72.50/hr

Integrate and fine-tune Large Language Models (LLMs) and other AI/ML models into enterprise applications. Develop and implement strategies for model deployment, inference, and monitoring, with an ...

Integrate and fine-tune Large Language Models (LLMs) and other AI/ML models into enterprise applications. Develop and implement strategies for model deployment, inference, and monitoring, with an ...

AI Solution Architect

Tempe, AZ · On-site

$60.25 - $79.50/hr

This individual will operate at the intersection of architecture, AI platform engineering, ML ... Real-time inference pipelines * Ensure architectural alignment with: * Cloud strategy * Enterprise ...

Experience serving ML models at scale using Triton Inference Server, TorchServe, Ray Serve, or similar. * Contributions to open-source ML projects or published research (papers, patents, or technical ...

Experience serving ML models at scale using Triton Inference Server, TorchServe, Ray Serve, or similar. * Contributions to open-source ML projects or published research (papers, patents, or technical ...

Data Scientist II

Phoenix, AZ · On-site

$120 - $160/hr

... inference techniques (uplift modeling, difference-in-differences, synthetic controls, instrumental variables) where randomized experiments aren't feasible * Translate business problems into ML ...

... inference techniques (uplift modeling, difference-in-differences, synthetic controls, instrumental variables) where randomized experiments aren't feasible * Translate business problems into ML ...

Data Scientist II

Phoenix, AZ · On-site

$120 - $170/hr

... inference techniques (uplift modeling, difference-in-differences, synthetic controls, instrumental variables) where randomized experiments aren't feasible * Translate business problems into ML ...

next page

Showing results 1-20

Ml Inference information

What is ML inference?

ML inference refers to the process of using a trained machine learning model to make predictions or decisions based on new data. After a model has been trained on historical data, inference is the phase where that model is deployed and used in real-world applications, such as recognizing speech, detecting objects in images, or recommending products. The focus in ML inference is on speed, efficiency, and scalability to ensure quick predictions, often in real time. This process is critical for practical applications like mobile apps, web services, and embedded systems. Optimizing inference involves reducing latency, memory usage, and computational requirements.

What are the key skills and qualifications needed to thrive in ML inference?

To thrive in ML Inference, you need a solid background in machine learning principles, programming (Python or C++), and experience with deploying models at scale, often supported by a degree in computer science or a related field. Familiarity with frameworks and tools such as TensorFlow, PyTorch, ONNX, and cloud platforms like AWS SageMaker or Google AI Platform is typically required. Strong problem-solving skills, attention to detail, and effective communication are crucial soft skills for collaborating with multidisciplinary teams and optimizing model performance. These skills ensure efficient, scalable, and reliable deployment of machine learning solutions in real-world applications.

What are some common challenges faced by ML inference engineers when deploying models to production?

ML Inference Engineers often encounter challenges such as optimizing model latency and throughput to meet production requirements, ensuring compatibility with diverse hardware environments, and managing model versioning and updates without disrupting service. Additionally, balancing resource utilization and inference accuracy while monitoring real-time performance metrics is crucial. Collaboration with data scientists, DevOps, and software engineers is typically essential to streamline deployment and maintain robust, scalable inference pipelines.

What is the difference between Ml Inference vs Data Scientist?

AspectML InferenceData Scientist
Required CredentialsKnowledge of machine learning models, programming skillsDegree in data science, statistics, or related fields
Work EnvironmentDeploying models in production, real-time data processingData analysis, model development, research
Industry UsageAI product deployment, software companiesResearch institutions, tech firms, consulting

ML Inference focuses on deploying trained models to make predictions on new data, often in real-time. Data Scientists develop and analyze models, working primarily in research and development. While both roles require understanding of machine learning, ML Inference emphasizes deployment and operationalization, whereas Data Scientists focus on model creation and analysis.

What job categories do people searching Ml Inference jobs in Arizona look for?

The top searched job categories for Ml Inference jobs in Arizona are:

What cities in Arizona are hiring for Ml Inference jobs?

Cities in Arizona with the most Ml Inference job openings:

IAM Engineer - Phoenix, Az

Motion Recruitment Partners, LLC

Phoenix, AZ • On-site

Other

Medical, Dental, Vision, Retirement, PTO

Posted 15 days ago


Job description


A leading financial services firm in Chandler, AZ is seeking a Senior Machine Learning Engineer to join their technology innovation team. In this position, you'll be at the forefront of building state-of-the-art machine learning solutions, taking ownership of projects from the design phase through to on-device deployment.
You'll architect and support robust pipelines for sensor data analysis, shaping ML systems that deliver real-time inference and anomaly detection on resource-constrained hardware. Collaborating closely with hardware, firmware, and platform teams, you'll be responsible for seamlessly integrating and validating intelligent features within embedded environments. This role will involve building and automating MLOps solutions, overseeing experiment management, and ensuring scalable and compliant machine learning lifecycle practices suited to regulatory demands.
Key Responsibilities
  • Lead the implementation of sensor data pipelines-from requirements through real-world integration
  • Design, train, optimize, and deploy ML models specifically for edge or embedded platforms, focusing on performance and reliability
  • Combine deep experience in Python with one or more frameworks (PyTorch, TensorFlow) to deliver production-ready models
  • Partner with engineering and product teams to embed, monitor, and validate ML inference in device applications
  • Develop infrastructure and automations for experiment tracking, continuous training, evaluation, and secure model deployment
  • Maintain full documentation for all aspects of the ML lifecycle to enable audit, regulatory compliance, and operational excellence
Required Skills & Experience
  • Extensive experience working with sensor and streaming data in production settings
  • Expertise in Python and at least one major ML library (PyTorch, TensorFlow, or alternatives)
  • Practical experience deploying models on edge systems (e.g., TensorRT, ONNX, TFLite, or similar technologies)
  • Solid engineering background with proficiency in C or C++ for embedded systems collaboration
  • Familiarity with MLOps and model versioning, preferably in a regulated environment
Preferred Qualifications
  • 5+ years hands-on ML development, with significant experience in real-time or embedded applications
  • Background working on medical, wearable, or robotics platforms highly desirable
  • Demonstrated success working in cross-functional teams with hardware/firmware integration
  • Experience in high-growth companies or small teams
  • Bachelor's or higher in Computer Science, Engineering, or a related field
Daily Activities
  • This role is hands-on and project-focused, where you'll design new ML workflows, transform raw data, and continually iterate on model performance
  • You'll support firmware engineers by ensuring seamless embedding and runtime validation of ML features
  • Work with DevOps and MLOps teams to automate deployments and create scalable, reliable ML solutions
  • Troubleshoot model issues, and refine algorithms for speed, accuracy, and compliance
Compensation & Benefits
Eligible for performance-based bonus or commission
Comprehensive benefits package including:
  • Health, dental, and vision insurance
  • Paid holidays and vacation
  • 401(k) retirement plan with company match (if applicable)
Applicants must have current authorization to work in the US on a full-time basis, now and in the future