The Machine Learning Engineer will be responsible for architecting and scaling quantization pipelines for multi-modal foundation models, optimizing inference latency and power consumption for various ...
The Machine Learning Engineer will be responsible for architecting and scaling quantization pipelines for multi-modal foundation models, optimizing inference latency and power consumption for various ...
Machine Learning Engineer
Palo Alto, CA · On-site
As a Machine Learning Engineer, you will play a central role in translating cutting-edge machine ... Hands on experience with quantization techniques (AWQ, GPTQ, FP8/GGUF)
Machine Learning Engineer
Palo Alto, CA · On-site
As a Machine Learning Engineer, you will play a central role in translating cutting-edge machine ... Hands on experience with quantization techniques (AWQ, GPTQ, FP8/GGUF)
Improve inference efficiency and model compression techniques, including quantization, pruning, and ... Engineering, Machine Learning, or related fields. * Must have prior experience managing a team ...
Improve inference efficiency and model compression techniques, including quantization, pruning, and ... Engineering, Machine Learning, or related fields. * Must have prior experience managing a team ...
Senior Machine Learning Engineer
San Jose, CA · On-site
$122K - $168K/yr
... quantization, pruning, and knowledge distillation. • Collaborate with cross-functional teams to ... Electrical Engineering, Machine Learning, or related fields. • Must have prior experience ...
Senior Machine Learning Engineer
San Jose, CA · On-site
$122K - $168K/yr
... quantization, pruning, and knowledge distillation. • Collaborate with cross-functional teams to ... Electrical Engineering, Machine Learning, or related fields. • Must have prior experience ...
Optimize models for production deployment, including ONNX / TensorRT / quantization / inference ... Electrical Engineering, Robotics, Computer Vision, Machine Learning, or a related field. * 3-5 ...
Optimize models for production deployment, including ONNX / TensorRT / quantization / inference ... Electrical Engineering, Robotics, Computer Vision, Machine Learning, or a related field. * 3-5 ...
Optimize models for production deployment, including ONNX / TensorRT / quantization / inference ... Electrical Engineering, Robotics, Computer Vision, Machine Learning, or a related field. * 3-5 ...
Optimize models for production deployment, including ONNX / TensorRT / quantization / inference ... Electrical Engineering, Robotics, Computer Vision, Machine Learning, or a related field. * 3-5 ...
With this mission, we are looking for passionated machine learning engineers of all levels who will ... Work on efficient foundation models (e.g. small LLMs, weight sharing, model quantization, etc ...
With this mission, we are looking for passionated machine learning engineers of all levels who will ... Work on efficient foundation models (e.g. small LLMs, weight sharing, model quantization, etc ...
... Level Machine Learning Engineer to develop and optimize machine learning models for edge AI ... quantization, pruning, and knowledge distillation. • Collaborate with cross-functional teams to ...
... Level Machine Learning Engineer to develop and optimize machine learning models for edge AI ... quantization, pruning, and knowledge distillation. • Collaborate with cross-functional teams to ...
Machine Learning Engineer Location: San Francisco, CA, USA (Hybrid/Remote) Job Type: Full-Time About the Role We are seeking an innovative Machine Learning Engineer to design, develop, and deploy ...
Machine Learning Engineer Location: San Francisco, CA, USA (Hybrid/Remote) Job Type: Full-Time About the Role We are seeking an innovative Machine Learning Engineer to design, develop, and deploy ...
Machine Learning Engineer
Pleasanton, CA · On-site
Machine Learning Engineer | Pleasanton, California, United States Machine Learning Engineer (Azure Focus) - Remote (PST) [About the Role] Join GAP as a Machine Learning Engineer and lead the ...
New
Machine Learning Engineer
Pleasanton, CA · On-site
Machine Learning Engineer | Pleasanton, California, United States Machine Learning Engineer (Azure Focus) - Remote (PST) [About the Role] Join GAP as a Machine Learning Engineer and lead the ...
New
Principal Machine Learning Engineer, Artificial Intelligence (AI) Required, Work From Home
San Francisco, CA · Remote
Principal Machine Learning Engineer, Artificial Intelligence (AI) Required, Work From Home As a ... quantization, and mixed precision. - Comfort owning ambiguous, zero-to-one ML systems end-to-end ...
Principal Machine Learning Engineer, Artificial Intelligence (AI) Required, Work From Home
San Francisco, CA · Remote
Principal Machine Learning Engineer, Artificial Intelligence (AI) Required, Work From Home As a ... quantization, and mixed precision. - Comfort owning ambiguous, zero-to-one ML systems end-to-end ...
Principal Machine Learning Engineer, Artificial Intelligence (AI) Required, Work From Home
San Francisco, CA · Remote
Principal Machine Learning Engineer, Artificial Intelligence (AI) Required, Work From Home As a ... quantization, and mixed precision. - Comfort owning ambiguous, zero-to-one ML systems end-to-end ...
Quick apply
Principal Machine Learning Engineer, Artificial Intelligence (AI) Required, Work From Home
San Francisco, CA · Remote
Principal Machine Learning Engineer, Artificial Intelligence (AI) Required, Work From Home As a ... quantization, and mixed precision. - Comfort owning ambiguous, zero-to-one ML systems end-to-end ...
Principal Machine Learning Engineer, Artificial Intelligence (AI) Required, Work From Home
San Francisco, CA · On-site +1
Company Description Principal Machine Learning Engineer, Artificial Intelligence (AI) Required ... quantization, and mixed precision. - Comfort owning ambiguous, zero-to-one ML systems end-to-end ...
Principal Machine Learning Engineer, Artificial Intelligence (AI) Required, Work From Home
San Francisco, CA · On-site +1
Company Description Principal Machine Learning Engineer, Artificial Intelligence (AI) Required ... quantization, and mixed precision. - Comfort owning ambiguous, zero-to-one ML systems end-to-end ...
They are seeking a Machine Learning Engineer to fine-tune pre-trained LLMs for use in Humanoid ... quantization, etc.) that can be deployed locally in Robots and Cars. • Develop and maintain the ...
They are seeking a Machine Learning Engineer to fine-tune pre-trained LLMs for use in Humanoid ... quantization, etc.) that can be deployed locally in Robots and Cars. • Develop and maintain the ...
They are seeking a Machine Learning Engineer to fine-tune pre-trained large language models (LLMs ... quantization, etc.) that can be deployed locally in Robots and Cars. • Develop and maintain the ...
They are seeking a Machine Learning Engineer to fine-tune pre-trained large language models (LLMs ... quantization, etc.) that can be deployed locally in Robots and Cars. • Develop and maintain the ...
Senior Machine Learning Engineer
$200K - $280K/yr
Improve inference efficiency and model compression techniques, including quantization, pruning, and ... Engineering, Machine Learning, or related fields. * Must have prior experience managing a team ...
Senior Machine Learning Engineer
$200K - $280K/yr
Improve inference efficiency and model compression techniques, including quantization, pruning, and ... Engineering, Machine Learning, or related fields. * Must have prior experience managing a team ...
Senior Machine Learning Engineer
San Jose, CA · On-site
$200K - $280K/yr
Improve inference efficiency and model compression techniques, including quantization, pruning, and ... Engineering, Machine Learning, or related fields. * Must have prior experience managing a team ...
Senior Machine Learning Engineer
San Jose, CA · On-site
$200K - $280K/yr
Improve inference efficiency and model compression techniques, including quantization, pruning, and ... Engineering, Machine Learning, or related fields. * Must have prior experience managing a team ...
Machine Learning Engineer
Santa Clara, CA · On-site
Job Description:Machine Learning Engineer The Role We are looking for a Machine Learning Engineer to join our core team building scalable ML systems for real-world perception and embodied ...
Quick apply
Machine Learning Engineer
Santa Clara, CA · On-site
Job Description:Machine Learning Engineer The Role We are looking for a Machine Learning Engineer to join our core team building scalable ML systems for real-world perception and embodied ...
Machine Learning Engineer
San Francisco, CA · On-site
$160K - $220K/yr
Optimize inference, batching, and quantization on GPU * Productionize models with clear SLAs and ... Solid engineering discipline and metrics focus Nice to have * Triton Inference Server, TensorRT ...
Machine Learning Engineer
San Francisco, CA · On-site
$160K - $220K/yr
Optimize inference, batching, and quantization on GPU * Productionize models with clear SLAs and ... Solid engineering discipline and metrics focus Nice to have * Triton Inference Server, TensorRT ...
Machine Learning Engineer
Fremont, CA · On-site
Machine Learning Engineer Location: Fremont, CA (Local) Onsite interview Duration: 12+ Mos H1B Only h1 candidate About the Role: Our direct client is hiring a Machine Learning Engineer for their ...
Quick apply
Machine Learning Engineer
Fremont, CA · On-site
Machine Learning Engineer Location: Fremont, CA (Local) Onsite interview Duration: 12+ Mos H1B Only h1 candidate About the Role: Our direct client is hiring a Machine Learning Engineer for their ...
Machine Learning Engineer Quantization information
See Sunnyvale, CA salary details
$37K - $54.3K
1% of jobs
$54.3K - $71.5K
1% of jobs
$71.5K - $88.8K
5% of jobs
$88.8K - $106.1K
6% of jobs
$120.4K is the 25th percentile. Wages below this are outliers.
$106.1K - $123.4K
14% of jobs
$123.4K - $140.7K
14% of jobs
The median wage is $149.3K / yr.
$140.7K - $158K
18% of jobs
$158K - $175.2K
14% of jobs
$178.8K is the 75th percentile. Wages above this are outliers.
$175.2K - $192.5K
12% of jobs
$192.5K - $209.8K
11% of jobs
$209.8K - $227.1K
5% of jobs
$37K
$151.1K
$227.1K
How much do machine learning engineer quantization jobs pay per year?
What are some common challenges Machine Learning Engineers face when implementing quantization techniques in production models?
What are the key skills and qualifications needed to thrive as a Machine Learning Engineer Quantization, and why are they important?
What does a Machine Learning Engineer Quantization do?
What is the difference between Machine Learning Engineer Quantization vs Data Scientist?
| Aspect | Machine Learning Engineer Quantization | Data Scientist |
|---|---|---|
| Required Credentials | Bachelor's or master's in CS, ML, or related; certifications in ML or AI | Bachelor's or master's in statistics, CS, or related; certifications in data analysis or statistics |
| Work Environment | Developing optimized ML models, deploying quantized models for efficiency | Analyzing data, building predictive models, interpreting results |
| Industry Usage | Tech companies, AI hardware firms, embedded systems | Finance, healthcare, marketing, research institutions |
Machine Learning Engineer Quantization focuses on optimizing ML models for deployment efficiency, often working closely with hardware and software teams. Data Scientists analyze data and build models for insights. While both roles require ML knowledge, quantization engineers specialize in model compression techniques, whereas data scientists focus on data analysis and interpretation.
- Artificial Intelligence Machine Learning Engineer
- Machine Learning Engineer Hybrid
- Machine Learning Engineer Software Engineer
- Physics Based Machine Learning
- Machine Learning Engineer Two
- Seasonal Medical Imaging Machine Learning
- Junior Machine Learning
- Machine Learning Biomedical Engineer
- Machine Learning Engineer Opt
- Contract Apple Machine Learning Engineer
Full-time
Re-posted 10 days ago
Tesla rating
8.5
Based on 679 frontline employees who took The Breakroom Quiz
1st of 44 rated automakers
Job description
Tesla is a leading company in the field of AI, focusing on building foundational models for real-world autonomy. The Machine Learning Engineer will be responsible for architecting and scaling quantization pipelines for multi-modal foundation models, optimizing inference latency and power consumption for various applications including self-driving cars and robots.
Responsibilities:
• Architect and scale quantization pipelines (both Post-Training Quantization and Quantization-Aware Training) for massive multi-modal foundation models that fuse vision, prediction, and decision-making. You will optimize inference latency, memory bandwidth utilization, and power consumption for self-driving cars, Optimus robots, and digital agents operating at enterprise scale
• Innovate quantization-aware-training recipes and algorithms that tackle complex optimization challenges inherent to low-precision training
• Push the limits of low-precision AI: Research and implement advanced low-bitweight post-training quantization techniques to address hard algorithmic problems such as activation outlier mitigation, KV cache compression, and optimal layer-wise bit-allocation while strictly maintaining model accuracy
• Collaborate closely with AI compiler, inference engine, and silicon teams to ensure models are architected to maximally utilize underlying hardware capabilities by co-designing quantization-friendly architectures, hardware-aware sparsity patterns, and mixed low-precision kernels
• Collaborate across perception, planning, robotics, digital agents, and infrastructure teams to move models from research to fleet-wide, robot-wide, and enterprise-wide deployment
Qualifications:
Required:
• Degree or equivalent experience in Computer Science, Machine Learning, Robotics, Computer Vision, or related quantitative field
• 2+ years of hands-on experience training, optimizing, and deploying large-scale quantized deep learning models
• Strong technical understanding of the challenges inherent to quantizing large transformer architectures, including mitigating massive activation outliers, KV cache quantization, and maintaining the numerical stability of attention mechanisms at low precision
• Deep expertise in the theory and low-level implementation of modern quantization algorithms (e.g., GPTQ, AWQ, SmoothQuant, OmniQuant)
• Experience with low-level numerics and emerging data formats (e.g., FP8, INT4, W4A8, W8A8, micro-scaling/MX formats) and their trade-offs regarding latency, memory bandwidth, and model fidelity
• Rigorous understanding of computer architecture and the roofline model. Familiarity with how to optimize for memory hierarchies, minimize SRAM/DRAM data movement, and efficiently map quantized GEMMs and memory-bound operators to custom silicon
• Proficiency in writing custom CUDA/Triton kernels, implementing custom autograd functions (e.g., Straight-Through Estimators), and manipulating PyTorch computational graphs (e.g., FX tracing, torch.compile)
• Strong software engineering skills — clean, production-grade Python/C++ code that ships reliably at scale
• Proven ability to turn cutting-edge research into robust, real-world systems that improve safety, capability, efficiency, or digital productivity
• Passion for Tesla’s mission and excitement about deploying AI that moves both the physical and digital worlds forward
Company:
Tesla is an electric vehicle and clean energy company that provides electric cars, solar, and renewable energy solutions. Founded in 2003, the company is headquartered in Austin, USA, with a team of 10001+ employees. The company is currently Late Stage.