1

Machine Learning Engineer Quantization Jobs in Phoenix, AZ

Design, develop, train, fine-tune, and evaluate machine learning models using PyTorch and ... Optimize model inference for latency and throughput; implement quantization, batching, and caching ...

Design, develop, train, fine-tune, and evaluate machine learning models using PyTorch and ... Optimize model inference for latency and throughput; implement quantization, batching, and caching ...

Senior Machine Learning Scientist

Scottsdale, AZ · On-site

$92K - $125K/yr

What You'll Do Location: any cities with Axon Engineering Hub in US, Vietnam, EU (see * US: Seattle ... in computer vision, machine learning, and deep learning, MLLMs, GenAI and integrate relevant ...

Core AI / Machine Learning * Deep expertise in : Machine learning fundamentals (supervised ... Edge AI or model compression/quantization * AI safety research and explainability techniques

Core AI / Machine Learning * Deep expertise in : Machine learning fundamentals (supervised ... Edge AI or model compression/quantization * AI safety research and explainability techniques

Sr. Advanced AI Software Engineer

Phoenix, AZ · On-site

$115K - $152K/yr

Core AI / Machine Learning * Deep expertise in : Machine learning fundamentals (supervised ... Edge AI or model compression/quantization * AI safety research and explainability techniques

Experience implementing and supporting endtoend Machine Learning workflows and patterns * Expert level programming skills in Python and experience with Data Science and ML packages and frameworks

AI/ML Engineer II

Phoenix, AZ · On-site +1

$113K - $136K/yr

Work with cross-functional team to contribute to machine learning projects throughout the machine learning lifecycle to include analysis, solution design, data pipeline engineering, testing ...

Work You'll Do As a Senior AI Engineer, you'll work cross-functionally with data scientists, machine learning engineers, project managers, and industry experts to develop robust AI infrastructure and ...

Research Engineer

Phoenix, AZ · On-site +1

$122K - $215K/yr

Qualifications: - Bachelor's in computer science, engineering, machine learning, or a related technical discipline. - Experience working on applied research projects. - Passion for taking research ...

Research Engineer

Phoenix, AZ · On-site +1

$122K - $215K/yr

Qualifications: - Bachelor's in computer science, engineering, machine learning, or a related technical discipline. - Experience working on applied research projects. - Passion for taking research ...

Master's degree + 2 years working experience in machine learning * Proficiency in at least one programming language such as Java, Python * Proficiency in big data, the use of frameworks related to ...

Master's degree + 2 years working experience in machine learning * Proficiency in at least one programming language such as Java, Python * Proficiency in big data, the use of frameworks related to ...

... machine learning, computer vision, and self-driving technologies, and apply insights from the ... Python programming with a focus on writing high-quality, well-structured, and tested code ...

... machine learning, computer vision, and self-driving technologies, and apply insights from the ... Python programming with a focus on writing high-quality, well-structured, and tested code ...

next page

Showing results 1-20

Machine Learning Engineer Quantization information

See Phoenix, AZ salary details

$31.3K

$127.9K

$192.1K

How much do machine learning engineer quantization jobs pay per year?

As of Jul 22, 2026, the average yearly pay for machine learning engineer quantization in Phoenix, AZ is $127,856.00, according to ZipRecruiter salary data. Most workers in this role earn between $100,800.00 and $153,900.00 per year, depending on experience, location, and employer.

What are some common challenges Machine Learning Engineers face when implementing quantization techniques in production models?

Machine Learning Engineers working on quantization often encounter challenges such as balancing reduced model size and computational efficiency with maintaining acceptable accuracy levels. Adapting quantization methods to different hardware platforms can also require significant testing and optimization. Additionally, engineers must frequently address compatibility issues with existing deployment pipelines and ensure that quantization-aware training is properly integrated to minimize performance degradation. Collaboration with hardware and software teams is essential to streamline deployment and achieve optimal results.

What are the key skills and qualifications needed to thrive as a Machine Learning Engineer Quantization, and why are they important?

To thrive as a Machine Learning Engineer Quantization, you need a solid background in machine learning, deep learning, and computer science, typically supported by a degree in a related field. Familiarity with quantization techniques, frameworks such as TensorFlow Lite or PyTorch, and experience with hardware accelerators are crucial. Strong problem-solving skills, attention to detail, and effective collaboration set top performers apart. These capabilities are vital for efficiently deploying high-performing models on resource-constrained devices and ensuring scalable, real-world AI solutions.

What does a Machine Learning Engineer Quantization do?

A Machine Learning Engineer specializing in quantization focuses on optimizing machine learning models by reducing their size and computational requirements without significantly sacrificing accuracy. This involves converting model parameters and computations from high-precision formats (like 32-bit floating point) to lower-precision formats (such as 8-bit integers). Quantization enables faster inference, lower memory usage, and allows models to run efficiently on edge devices and mobile platforms. These engineers work closely with data scientists and hardware teams to implement, test, and validate quantized models in production environments.

What is the difference between Machine Learning Engineer Quantization vs Data Scientist?

AspectMachine Learning Engineer QuantizationData Scientist
Required CredentialsBachelor's or master's in CS, ML, or related; certifications in ML or AIBachelor's or master's in statistics, CS, or related; certifications in data analysis or statistics
Work EnvironmentDeveloping optimized ML models, deploying quantized models for efficiencyAnalyzing data, building predictive models, interpreting results
Industry UsageTech companies, AI hardware firms, embedded systemsFinance, healthcare, marketing, research institutions

Machine Learning Engineer Quantization focuses on optimizing ML models for deployment efficiency, often working closely with hardware and software teams. Data Scientists analyze data and build models for insights. While both roles require ML knowledge, quantization engineers specialize in model compression techniques, whereas data scientists focus on data analysis and interpretation.

What are popular job titles related to Machine Learning Engineer Quantization jobs in Phoenix, AZ? For Machine Learning Engineer Quantization jobs in Phoenix, AZ, the most frequently searched job titles are:
AI Engineer

Full-time

Medical, Dental, Vision, Retirement, PTO

Posted 8 days ago


Entertainment Partners rating

8.2

Company rating: 8.2 out of 10

Based on 17 frontline employees who took The Breakroom Quiz


Job description


At Entertainment Partners and Central Casting, we are committed to creating an environment where every employee is seen, where ideas, thoughts and perspectives are shared openly, and where fearless innovation is encouraged. Weaving diversity, equity, and inclusion into who we are will drive our competitiveness by encouraging creativity and enhanced decision making.
We help to power Oscar-winning films, Emmy-winning shows, and Clio-winning commercials. Feel the satisfaction of doing work that directly impacts the most exciting industry in the world. EP is poised to redefine and evolve the back-office processes of the entertainment community with security at the core of what we do.
Are you looking for the next opportunity to revolutionize an industry? If so....
Entertainment Partners (EP) is seeking a Senior Software Engineer specializing in AI and Machine Learning to join our AI Services organization. This role sits at the intersection of applied ML engineering, LLM product development, and production-grade system design. The AI Senior Software Engineer is responsible for building, training, evaluating, and deploying AI/ML models and agentic systems that power EP's intelligent product suite - including Rosey Intelligence, Project Florence, and EP Answers. The ideal candidate brings deep hands-on expertise in PyTorch, transformer architectures, and the full ML lifecycle, combined with the software engineering discipline required to ship reliable AI products at scale in a production entertainment technology environment.
KEY RESPONSIBILITIES
In addition to the following, other duties may be assigned to meet business needs.
AI / ML Engineering
  • Design, develop, train, fine-tune, and evaluate machine learning models using PyTorch and associated ecosystem libraries (torchvision, torchaudio, torch.nn, torch.optim).
  • Build and maintain ML training pipelines, experiment tracking workflows, and model evaluation frameworks.
  • Implement transformer-based models and large language model (LLM) integrations for production use cases including NLP, information extraction, classification, and generation.
  • Apply parameter-efficient fine-tuning techniques (LoRA, QLoRA, PEFT) to adapt foundation models for EP-specific domains (payroll, residuals, production management).
  • Design and implement RAG (Retrieval-Augmented Generation) architectures using vector databases (pgvector, Pinecone, Weaviate) and semantic search pipelines.
  • Optimize model inference for latency and throughput; implement quantization, batching, and caching strategies for production serving.
  • Develop and maintain AI evaluation frameworks - including automated evals as unit tests - to ensure model behavior is reliable, safe, and production-grade.

LLM Integration & Agentic Systems
  • Design and implement LLM-powered agentic workflows using LangChain, LangGraph, and EP's internal MCP (Model Context Protocol) server architecture.
  • Build multi-step reasoning pipelines, tool-calling agents, and autonomous task execution systems that integrate with EP's enterprise data and product APIs.
  • Implement prompt engineering strategies, few-shot templates, chain-of-thought scaffolding, and structured output validation.
  • Apply and maintain EP's AI quality engineering (QE) standards including failure taxonomy, runtime guardrails, and evidence-driven release gates.
  • Contribute to EP's Enterprise Context Engine - the governed, zero-data-retention AI context layer exposed via MCP to Tabnine Agent and Claude Code.
  • MLOps & Production Engineering
  • Build and maintain MLOps infrastructure for model training, experiment tracking (MLflow, Weights & Biases), versioning, and deployment.
  • Containerize and deploy ML services using Docker and Kubernetes; integrate with CI/CD pipelines (GitHub Actions, Azure DevOps).
  • Monitor model performance in production; implement drift detection, feedback loops, and automated retraining triggers.
  • Ensure AI systems meet EP's security, privacy, and compliance requirements including data minimization and access control for sensitive payroll data.
  • Collaborate with the data engineering team to design and maintain feature stores, data pipelines, and training data infrastructure.

Collaboration & Technical Leadership
  • Partner with the Chief Architect AI & Data and CAIO to define AI architecture patterns and best practices for the EP engineering organization.
  • Collaborate with product managers, UX designers, and full stack engineers to translate AI capabilities into well-designed product features.
  • Conduct code reviews for AI/ML code with a focus on reproducibility, correctness, and production readiness.
  • Mentor engineers across the organization in AI engineering fundamentals, LLM integration patterns, and responsible AI practices.
  • Stay current with the rapidly evolving AI/ML landscape; evaluate new models, frameworks, and techniques for potential application at EP.
  • Contribute to EP's PE AI Maturity Scorecard (S1-S3) by advancing the organization's AI capability maturity.
  • Represent EP's AI engineering practices in Architecture Review Board discussions.

JOB REQUIREMENTS / QUALIFICATIONS NEEDED
Minimum qualifications:
  • Bachelor's or Master's degree in Computer Science, Machine Learning, Statistics, Mathematics, or a related quantitative field.
  • 6-10+ years of professional software engineering experience, with a minimum of 3+ years focused on ML/AI engineering in production environments.
  • Expert-level proficiency in Python; deep familiarity with the Python ML/AI ecosystem.
  • Hands-on production experience with PyTorch - model definition (nn.Module), custom training loops, autograd, GPU acceleration (CUDA), and model serialization (TorchScript, ONNX).
  • Experience with Hugging Face Transformers, Datasets, and PEFT libraries; ability to fine-tune and adapt foundation models.
  • Demonstrated experience building RAG pipelines, including chunking strategies, embedding models, vector store selection, and retrieval evaluation.
  • Production experience integrating LLM APIs (OpenAI, Anthropic, open-source via vLLM/Ollama) and building reliable prompt engineering systems.
  • Experience with LangChain or LangGraph for multi-step agent and tool-calling workflows.
  • Strong understanding of ML fundamentals: supervised/unsupervised learning, loss functions, regularization, evaluation metrics, and statistical validation.
  • Experience with experiment tracking tools (MLflow, Weights & Biases, Comet) and reproducible ML workflows.
  • Working knowledge of containerization (Docker) and cloud ML services (AWS SageMaker, Azure ML, or OCI Data Science).
  • Experience with SQL and NoSQL databases; ability to design data pipelines for ML training and inference.

Preferred qualifications:
  • Experience with additional deep learning frameworks (TensorFlow, JAX) or framework interoperability (ONNX).
  • Familiarity with computer vision (torchvision, OpenCV) or speech/audio processing (torchaudio) domains.
  • Experience with model compression techniques: quantization (INT8, FP16, BF16), pruning, distillation.
  • Experience serving ML models at scale using Triton Inference Server, TorchServe, Ray Serve, or similar.
  • Contributions to open-source ML projects or published research (papers, patents, or technical blog posts).
  • Experience with responsible AI frameworks, bias evaluation, and AI governance practices.
  • Familiarity with MCP (Model Context Protocol) server development for exposing tools to AI agents.
  • Prior domain experience in payroll, fintech, media, or enterprise SaaS environments.
  • Experience with Kubernetes-based ML workload orchestration (Kubeflow, KFServing, or similar).
  • Hybrid work environment - Burbank, CA headquarters with flexible remote schedule.
  • On-call availability as needed for production AI system incidents and model deployment events.
  • Access to GPU-accelerated compute environments (cloud-based) for model training workloads.
  • Sitting for extended periods of time at a computer workstation.
  • Dexterity of hands and fingers to operate a computer keyboard and mouse.
  • Occasional participation in early-morning or evening sessions to coordinate with distributed teams or international partners.

Other benefits and perks included are:
  • Health, Dental, and Vision options
  • 401(k) retirement savings plan and company match
  • Paid holidays, vacation time, and sick time
  • Participation in company equity plans
  • Employee Assistance Program, mental health and wellness programs
  • Training and development
  • Annual bonus and merit reviews

The salary range for this position in $140,000 to $180,000 and will be commensurate with experience related to the position.
Entertainment Partners seeks to employ the most qualified individuals from the available workforce and to provide equal employment opportunity for all persons. Our policy prohibits unlawful discrimination based on race, color, religion, religious creed, sex, gender identity/expression, age, pregnancy, citizenship status, marital status, national origin or ancestry, physical or mental disability (whether perceived or actual), medical condition (cancer-related or genetic characteristics-related), sexual orientation, veteran status, medical/family care leave status or any other consideration made unlawful by applicable federal, state, or local laws. Qualified applicants with arrest or conviction records will be considered for employment in accordance with the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act.
Equal opportunity extends to all aspects of the employment relationship, including recruiting, hiring, transfers, promotions, training, terminations, working conditions, compensation, benefits, and other terms and conditions of employment.

What Entertainment Partners employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom