1

Machine Learning Engineer Quantization Jobs in Sunnyvale, CA

As a Machine Learning Engineer, you will play a central role in translating cutting-edge machine ... Hands on experience with quantization techniques (AWQ, GPTQ, FP8/GGUF)

Optimize models for production deployment, including ONNX / TensorRT / quantization / inference ... Electrical Engineering, Robotics, Computer Vision, Machine Learning, or a related field. * 3-5 ...

Improve inference efficiency and model compression techniques, including quantization, pruning, and ... Engineering, Machine Learning, or related fields. * Must have prior experience managing a team ...

Machine Learning Engineer Location: Fremont, CA (Local) Onsite interview Duration: 12+ Mos H1B Only h1 candidate About the Role: Our direct client is hiring a Machine Learning Engineer for their ...

next page

Showing results 1-20

Machine Learning Engineer Quantization information

See Sunnyvale, CA salary details

$37K

$151.1K

$227.1K

How much do machine learning engineer quantization jobs pay per year?

As of Jul 30, 2026, the average yearly pay for machine learning engineer quantization in Sunnyvale, CA is $151,130.00, according to ZipRecruiter salary data. Most workers in this role earn between $119,100.00 and $181,900.00 per year, depending on experience, location, and employer.

What are some common challenges Machine Learning Engineers face when implementing quantization techniques in production models?

Machine Learning Engineers working on quantization often encounter challenges such as balancing reduced model size and computational efficiency with maintaining acceptable accuracy levels. Adapting quantization methods to different hardware platforms can also require significant testing and optimization. Additionally, engineers must frequently address compatibility issues with existing deployment pipelines and ensure that quantization-aware training is properly integrated to minimize performance degradation. Collaboration with hardware and software teams is essential to streamline deployment and achieve optimal results.

What are the key skills and qualifications needed to thrive as a Machine Learning Engineer Quantization, and why are they important?

To thrive as a Machine Learning Engineer Quantization, you need a solid background in machine learning, deep learning, and computer science, typically supported by a degree in a related field. Familiarity with quantization techniques, frameworks such as TensorFlow Lite or PyTorch, and experience with hardware accelerators are crucial. Strong problem-solving skills, attention to detail, and effective collaboration set top performers apart. These capabilities are vital for efficiently deploying high-performing models on resource-constrained devices and ensuring scalable, real-world AI solutions.

What does a Machine Learning Engineer Quantization do?

A Machine Learning Engineer specializing in quantization focuses on optimizing machine learning models by reducing their size and computational requirements without significantly sacrificing accuracy. This involves converting model parameters and computations from high-precision formats (like 32-bit floating point) to lower-precision formats (such as 8-bit integers). Quantization enables faster inference, lower memory usage, and allows models to run efficiently on edge devices and mobile platforms. These engineers work closely with data scientists and hardware teams to implement, test, and validate quantized models in production environments.

What is the difference between Machine Learning Engineer Quantization vs Data Scientist?

AspectMachine Learning Engineer QuantizationData Scientist
Required CredentialsBachelor's or master's in CS, ML, or related; certifications in ML or AIBachelor's or master's in statistics, CS, or related; certifications in data analysis or statistics
Work EnvironmentDeveloping optimized ML models, deploying quantized models for efficiencyAnalyzing data, building predictive models, interpreting results
Industry UsageTech companies, AI hardware firms, embedded systemsFinance, healthcare, marketing, research institutions

Machine Learning Engineer Quantization focuses on optimizing ML models for deployment efficiency, often working closely with hardware and software teams. Data Scientists analyze data and build models for insights. While both roles require ML knowledge, quantization engineers specialize in model compression techniques, whereas data scientists focus on data analysis and interpretation.

What are popular job titles related to Machine Learning Engineer Quantization jobs in Sunnyvale, CA? For Machine Learning Engineer Quantization jobs in Sunnyvale, CA, the most frequently searched job titles are:
What cities near Sunnyvale, CA are hiring for Machine Learning Engineer Quantization jobs? Cities near Sunnyvale, CA with the most Machine Learning Engineer Quantization job openings:

Machine Learning Engineer, Model Quantization, Tesla AI

Tesla

Palo Alto, CA • On-site

Full-time

Re-posted 10 days ago


Tesla rating

8.5

Company rating: 8.5 out of 10

Based on 679 frontline employees who took The Breakroom Quiz

1st of 44 rated automakers


Job description

Job Summary:
Tesla is a leading company in the field of AI, focusing on building foundational models for real-world autonomy. The Machine Learning Engineer will be responsible for architecting and scaling quantization pipelines for multi-modal foundation models, optimizing inference latency and power consumption for various applications including self-driving cars and robots.
Responsibilities:
• Architect and scale quantization pipelines (both Post-Training Quantization and Quantization-Aware Training) for massive multi-modal foundation models that fuse vision, prediction, and decision-making. You will optimize inference latency, memory bandwidth utilization, and power consumption for self-driving cars, Optimus robots, and digital agents operating at enterprise scale
• Innovate quantization-aware-training recipes and algorithms that tackle complex optimization challenges inherent to low-precision training
• Push the limits of low-precision AI: Research and implement advanced low-bitweight post-training quantization techniques to address hard algorithmic problems such as activation outlier mitigation, KV cache compression, and optimal layer-wise bit-allocation while strictly maintaining model accuracy
• Collaborate closely with AI compiler, inference engine, and silicon teams to ensure models are architected to maximally utilize underlying hardware capabilities by co-designing quantization-friendly architectures, hardware-aware sparsity patterns, and mixed low-precision kernels
• Collaborate across perception, planning, robotics, digital agents, and infrastructure teams to move models from research to fleet-wide, robot-wide, and enterprise-wide deployment
Qualifications:
Required:
• Degree or equivalent experience in Computer Science, Machine Learning, Robotics, Computer Vision, or related quantitative field
• 2+ years of hands-on experience training, optimizing, and deploying large-scale quantized deep learning models
• Strong technical understanding of the challenges inherent to quantizing large transformer architectures, including mitigating massive activation outliers, KV cache quantization, and maintaining the numerical stability of attention mechanisms at low precision
• Deep expertise in the theory and low-level implementation of modern quantization algorithms (e.g., GPTQ, AWQ, SmoothQuant, OmniQuant)
• Experience with low-level numerics and emerging data formats (e.g., FP8, INT4, W4A8, W8A8, micro-scaling/MX formats) and their trade-offs regarding latency, memory bandwidth, and model fidelity
• Rigorous understanding of computer architecture and the roofline model. Familiarity with how to optimize for memory hierarchies, minimize SRAM/DRAM data movement, and efficiently map quantized GEMMs and memory-bound operators to custom silicon
• Proficiency in writing custom CUDA/Triton kernels, implementing custom autograd functions (e.g., Straight-Through Estimators), and manipulating PyTorch computational graphs (e.g., FX tracing, torch.compile)
• Strong software engineering skills — clean, production-grade Python/C++ code that ships reliably at scale
• Proven ability to turn cutting-edge research into robust, real-world systems that improve safety, capability, efficiency, or digital productivity
• Passion for Tesla’s mission and excitement about deploying AI that moves both the physical and digital worlds forward
Company:
Tesla is an electric vehicle and clean energy company that provides electric cars, solar, and renewable energy solutions. Founded in 2003, the company is headquartered in Austin, USA, with a team of 10001+ employees. The company is currently Late Stage.

What Tesla employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom