1

Machine Learning Engineer Quantization Jobs in Utah

Senior Machine Learning Engineer

Sandy, UT ยท Hybrid

$99K - $136K/yr

As Senior Machine Learning Engineer, you will own the evaluation and optimization of speech ... LoRA/PEFT for speech models, inference optimization (quantization, SGLang/vLLM serving for audio ...

They are seeking a highly skilled Machine Learning Engineer to manage large datasets, optimize cloud-based computing resources, and train advanced machine-learning models to contribute to new ...

Senior Machine Learning Engineer

Sandy, UT ยท On-site

$113K - $150K/yr

As Senior Machine Learning Engineer, you will own the evaluation and optimization of speech ... LoRA/PEFT for speech models, inference optimization (quantization, SGLang/vLLM serving for audio ...

Job Summary As a Machine Learning Engineer II, you will lead the productization of AI/ML research ... Optimize ML model pipelines for inference efficiency using techniques such as quantization ...

Must-Have Skills 3+ years of ML engineering experience -- model training, fine-tuning, or post-training pipelines in research or production Strong Python and deep learning proficiency (PyTorch ...

Must-Have Skills 3+ years of ML engineering experience -- model training, fine-tuning, or post-training pipelines in research or production Strong Python and deep learning proficiency (PyTorch ...

Must-Have Skills 3+ years of ML engineering experience -- model training, fine-tuning, or post-training pipelines in research or production Strong Python and deep learning proficiency (PyTorch ...

Must-Have Skills 3+ years of ML engineering experience -- model training, fine-tuning, or post-training pipelines in research or production Strong Python and deep learning proficiency (PyTorch ...

As a Machine Learning Engineer, you will have the opportunity to collaborate closely with senior engineers and product leaders as part of your team. Together, you'll develop and enhance Instacart ...

$139K - $168K/yr

Our team of Machine Learning Engineers have high impact by advancing the current Machine Learning systems, building performant and reliable LLM applications and collaborating with our product team to ...

next page

Showing results 1-20

Machine Learning Engineer Quantization information

What are some common challenges machine learning engineers face when implementing quantization techniques in production models?

Machine Learning Engineers working on quantization often encounter challenges such as balancing reduced model size and computational efficiency with maintaining acceptable accuracy levels. Adapting quantization methods to different hardware platforms can also require significant testing and optimization. Additionally, engineers must frequently address compatibility issues with existing deployment pipelines and ensure that quantization-aware training is properly integrated to minimize performance degradation. Collaboration with hardware and software teams is essential to streamline deployment and achieve optimal results.

What are the key skills and qualifications needed to thrive as a machine learning engineer quantization, and why are they important?

To thrive as a Machine Learning Engineer Quantization, you need a solid background in machine learning, deep learning, and computer science, typically supported by a degree in a related field. Familiarity with quantization techniques, frameworks such as TensorFlow Lite or PyTorch, and experience with hardware accelerators are crucial. Strong problem-solving skills, attention to detail, and effective collaboration set top performers apart. These capabilities are vital for efficiently deploying high-performing models on resource-constrained devices and ensuring scalable, real-world AI solutions.

What does a machine learning engineer quantization do?

A Machine Learning Engineer specializing in quantization focuses on optimizing machine learning models by reducing their size and computational requirements without significantly sacrificing accuracy. This involves converting model parameters and computations from high-precision formats (like 32-bit floating point) to lower-precision formats (such as 8-bit integers). Quantization enables faster inference, lower memory usage, and allows models to run efficiently on edge devices and mobile platforms. These engineers work closely with data scientists and hardware teams to implement, test, and validate quantized models in production environments.

What is the difference between Machine Learning Engineer Quantization vs Data Scientist?

AspectMachine Learning Engineer QuantizationData Scientist
Required CredentialsBachelor's or master's in CS, ML, or related; certifications in ML or AIBachelor's or master's in statistics, CS, or related; certifications in data analysis or statistics
Work EnvironmentDeveloping optimized ML models, deploying quantized models for efficiencyAnalyzing data, building predictive models, interpreting results
Industry UsageTech companies, AI hardware firms, embedded systemsFinance, healthcare, marketing, research institutions

Machine Learning Engineer Quantization focuses on optimizing ML models for deployment efficiency, often working closely with hardware and software teams. Data Scientists analyze data and build models for insights. While both roles require ML knowledge, quantization engineers specialize in model compression techniques, whereas data scientists focus on data analysis and interpretation.

What cities in Utah are hiring for Machine Learning Engineer Quantization jobs? Cities in Utah with the most Machine Learning Engineer Quantization job openings:
Infographic showing various Machine Learning Engineer Quantization job openings in Utah as of August 2026, with employment types broken down into 1% As Needed, 74% Full Time, 22% Part Time, 1% Temporary, and 2% Contract. Highlights an 89% Physical, 2% Hybrid, and 9% Remote job distribution.

Senior Machine Learning Engineer

NICE

Sandy, UT โ€ข Hybrid

$99K - $136K/yr

Full-time

Posted 28 days ago


Job description

So, what's the role all about?

NiCE is looking for a Senior Machine Learning Engineer to join NiCE Labs Research (NLR), a team dedicated to model expertise and agent architecture for the Cognigy platform. As Senior Machine Learning Engineer, you will own the evaluation and optimization of speech-oriented AI models - covering real-time transcription and speech-to-speech systems across dozens of languages.

This role is primarily concerned with rigorous measurement: designing test suites, running comparative evaluations, and producing actionable recommendations on model selection and configuration.

The Senior Machine Learning Engineer monitors the rapidly evolving speech AI landscape to identify state-of-the-art transcription and speech-to-speech models for evaluation. You will design and maintain a speech-oriented test suite that covers quality, cost, and latency, and develop techniques to optimize model usage for operational deployment.

This role requires deep expertise in speech AI systems, strong quantitative skills, and the discipline to produce reliable, reproducible evaluation results.

How will you make an impact?

  • Design and maintain a speech-oriented test suite covering quality, cost, and latency across dozens of languages.
  • Monitor the industry for new state-of-the-art transcription and speech-to-speech models to evaluate.
  • Design and evaluate techniques to optimize speech model usage for operational deployment.
  • Produce clear, quantitative evaluation reports and model recommendations for technical and non-technical stakeholders.
  • Contribute to the broader model evaluation framework maintained by the NLR team.
  • Stay informed of advances in speech AI, including transcription, text-to-speech, and speech-to-speech technologies.

Have you got what it takes?

  • MS in computer science, electrical engineering, computational linguistics, or a related field with a focus on speech or audio processing.
  • Three or more years of hands-on experience with speech AI systems, including ASR, TTS, or speech-to-speech models.
  • Experience designing evaluation methodologies or test suites for AI systems.
  • Strong quantitative and analytical skills, with experience producing rigorous benchmark results.
  • LoRA/PEFT for speech models, inference optimization (quantization, SGLang/vLLM serving for audio, distillation), experience with at least one open-source TTS family
  • GPU cost modeling
  • Proficiency in Python and familiarity with speech processing libraries and tools.
  • Experience with cloud-based infrastructure (AWS, Azure, or GCP).
  • Ability to develop and maintain good working relationships with cross-functional teams.
  • Ability to clearly communicate and present to internal and external stakeholders.

You will have an advantage if you have:

  • Experience evaluating speech models across multiple languages.
  • Familiarity with multi-cloud deployment across AWS, Azure, and Google Cloud.
  • Experience with model optimization techniques for speech systems, such as latency reduction or cost optimization.
  • Exposure to contact center or conversational AI platforms.
  • Experience working on international, globe-spanning teams.

ย 

What's in it for you?

Join an ever-growing, market disrupting, global company where the teams - comprised of the best of the best - work in a fast-paced, collaborative, and creative environment! As the market leader, every day at NiCE is a chance to learn and grow, and there are endless internal career opportunities across multiple roles, disciplines, domains, and locations. If you are passionate, innovative, and excited to constantly raise the bar, you may just be our next NICEr!

Enjoy NiCE-FLEX!

At NiCE, we work according to the NiCE-FLEX hybrid model, which enables maximum flexibility: 2 days working from the office and 3 days of remote work, each week. Naturally, office days focus on face-to-face meetings, where teamwork and collaborative thinking generate innovation, new ideas, and a vibrant, interactive atmosphere.

ย 

Requisition ID: 11422

Reporting into: Director, Engineering, AI Research, NiCE Labs

Role Type: Individual Contributor