The Machine Learning Engineer will be responsible for architecting and scaling quantization pipelines for multi-modal foundation models, optimizing inference latency and power consumption for various ...
The Machine Learning Engineer will be responsible for architecting and scaling quantization pipelines for multi-modal foundation models, optimizing inference latency and power consumption for various ...
Machine Learning Engineer
Palo Alto, CA · On-site
As a Machine Learning Engineer, you will play a central role in translating cutting-edge machine ... Hands on experience with quantization techniques (AWQ, GPTQ, FP8/GGUF)
Machine Learning Engineer
Palo Alto, CA · On-site
As a Machine Learning Engineer, you will play a central role in translating cutting-edge machine ... Hands on experience with quantization techniques (AWQ, GPTQ, FP8/GGUF)
Improve inference efficiency and model compression techniques, including quantization, pruning, and ... Engineering, Machine Learning, or related fields. * Must have prior experience managing a team ...
Improve inference efficiency and model compression techniques, including quantization, pruning, and ... Engineering, Machine Learning, or related fields. * Must have prior experience managing a team ...
Senior Machine Learning Engineer
San Jose, CA · On-site
$122K - $168K/yr
... quantization, pruning, and knowledge distillation. • Collaborate with cross-functional teams to ... Electrical Engineering, Machine Learning, or related fields. • Must have prior experience ...
Senior Machine Learning Engineer
San Jose, CA · On-site
$122K - $168K/yr
... quantization, pruning, and knowledge distillation. • Collaborate with cross-functional teams to ... Electrical Engineering, Machine Learning, or related fields. • Must have prior experience ...
As a Machine Learning Engineer, you will shape the technical direction of the company by automating ... quantization. Qualifications : Required : • 3 to 5 years of industry experience in full-stack ...
As a Machine Learning Engineer, you will shape the technical direction of the company by automating ... quantization. Qualifications : Required : • 3 to 5 years of industry experience in full-stack ...
Optimize models for production deployment, including ONNX / TensorRT / quantization / inference ... Electrical Engineering, Robotics, Computer Vision, Machine Learning, or a related field. * 3-5 ...
Optimize models for production deployment, including ONNX / TensorRT / quantization / inference ... Electrical Engineering, Robotics, Computer Vision, Machine Learning, or a related field. * 3-5 ...
Optimize models for production deployment, including ONNX / TensorRT / quantization / inference ... Electrical Engineering, Robotics, Computer Vision, Machine Learning, or a related field. * 3-5 ...
Optimize models for production deployment, including ONNX / TensorRT / quantization / inference ... Electrical Engineering, Robotics, Computer Vision, Machine Learning, or a related field. * 3-5 ...
With this mission, we are looking for passionated machine learning engineers of all levels who will ... Work on efficient foundation models (e.g. small LLMs, weight sharing, model quantization, etc ...
With this mission, we are looking for passionated machine learning engineers of all levels who will ... Work on efficient foundation models (e.g. small LLMs, weight sharing, model quantization, etc ...
... Level Machine Learning Engineer to develop and optimize machine learning models for edge AI ... quantization, pruning, and knowledge distillation. • Collaborate with cross-functional teams to ...
... Level Machine Learning Engineer to develop and optimize machine learning models for edge AI ... quantization, pruning, and knowledge distillation. • Collaborate with cross-functional teams to ...
Senior Machine Learning Engineer, Proactive
Santa Clara, CA · On-site
$143K - $189K/yr
At Apple, machine learning powers experiences that anticipate what people need before they ask. We ... quantization, and low-latency inference. You'll partner with engineers, researchers, product ...
New
Senior Machine Learning Engineer, Proactive
Santa Clara, CA · On-site
$143K - $189K/yr
At Apple, machine learning powers experiences that anticipate what people need before they ask. We ... quantization, and low-latency inference. You'll partner with engineers, researchers, product ...
New
Machine Learning Engineer Location: San Francisco, CA, USA (Hybrid/Remote) Job Type: Full-Time About the Role We are seeking an innovative Machine Learning Engineer to design, develop, and deploy ...
Machine Learning Engineer Location: San Francisco, CA, USA (Hybrid/Remote) Job Type: Full-Time About the Role We are seeking an innovative Machine Learning Engineer to design, develop, and deploy ...
The Machine Learning Engineer will develop and optimize machine learning models to enhance healthcare operations and collaborate with cross-functional teams to integrate solutions into existing ...
The Machine Learning Engineer will develop and optimize machine learning models to enhance healthcare operations and collaborate with cross-functional teams to integrate solutions into existing ...
Principal Machine Learning Engineer, Artificial Intelligence (AI) Required, Work From Home
San Francisco, CA · Remote
Principal Machine Learning Engineer, Artificial Intelligence (AI) Required, Work From Home As a ... quantization, and mixed precision. - Comfort owning ambiguous, zero-to-one ML systems end-to-end ...
Quick apply
Principal Machine Learning Engineer, Artificial Intelligence (AI) Required, Work From Home
San Francisco, CA · Remote
Principal Machine Learning Engineer, Artificial Intelligence (AI) Required, Work From Home As a ... quantization, and mixed precision. - Comfort owning ambiguous, zero-to-one ML systems end-to-end ...
Principal Machine Learning Engineer, Artificial Intelligence (AI) Required, Work From Home
San Francisco, CA · Remote
Principal Machine Learning Engineer, Artificial Intelligence (AI) Required, Work From Home As a ... quantization, and mixed precision. - Comfort owning ambiguous, zero-to-one ML systems end-to-end ...
Principal Machine Learning Engineer, Artificial Intelligence (AI) Required, Work From Home
San Francisco, CA · Remote
Principal Machine Learning Engineer, Artificial Intelligence (AI) Required, Work From Home As a ... quantization, and mixed precision. - Comfort owning ambiguous, zero-to-one ML systems end-to-end ...
Machine Learning Engineer
Pleasanton, CA · On-site
Machine Learning Engineer | Pleasanton, California, United States Machine Learning Engineer (Azure Focus) - Remote (PST) [About the Role] Join GAP as a Machine Learning Engineer and lead the ...
Machine Learning Engineer
Pleasanton, CA · On-site
Machine Learning Engineer | Pleasanton, California, United States Machine Learning Engineer (Azure Focus) - Remote (PST) [About the Role] Join GAP as a Machine Learning Engineer and lead the ...
Principal Machine Learning Engineer, Artificial Intelligence (AI) Required, Work From Home
San Francisco, CA · On-site +1
Company Description Principal Machine Learning Engineer, Artificial Intelligence (AI) Required ... quantization, and mixed precision. - Comfort owning ambiguous, zero-to-one ML systems end-to-end ...
Principal Machine Learning Engineer, Artificial Intelligence (AI) Required, Work From Home
San Francisco, CA · On-site +1
Company Description Principal Machine Learning Engineer, Artificial Intelligence (AI) Required ... quantization, and mixed precision. - Comfort owning ambiguous, zero-to-one ML systems end-to-end ...
Machine Learning Engineer
Carlsbad, CA · On-site
Machine Learning Engineer Position: Full time Location: Carlsbad office About Us: NTENT provides a Platform-as-a-Service (PaaS), allowing industry partners to customize, localize and integrate search ...
Machine Learning Engineer
Carlsbad, CA · On-site
Machine Learning Engineer Position: Full time Location: Carlsbad office About Us: NTENT provides a Platform-as-a-Service (PaaS), allowing industry partners to customize, localize and integrate search ...
Machine Learning Engineer Position: Full time Location: Carlsbad office About Us: NTENT provides a Platform-as-a-Service (PaaS), allowing industry partners to customize, localize and integrate search ...
Machine Learning Engineer Position: Full time Location: Carlsbad office About Us: NTENT provides a Platform-as-a-Service (PaaS), allowing industry partners to customize, localize and integrate search ...
Machine Learning Engineer
Los Angeles, CA · On-site
They are seeking a Machine Learning Engineer to develop solutions that enhance performance in wireless communication systems, impacting products and services that utilize signal processing technology.
Machine Learning Engineer
Los Angeles, CA · On-site
They are seeking a Machine Learning Engineer to develop solutions that enhance performance in wireless communication systems, impacting products and services that utilize signal processing technology.
Machine Learning Engineer Full-time Responsibilities * Build, maintain, and improve efficient and reliable data mining and machine learning models. * Design, implement and tune machine learning ...
Machine Learning Engineer Full-time Responsibilities * Build, maintain, and improve efficient and reliable data mining and machine learning models. * Design, implement and tune machine learning ...
Machine Learning Engineer Quantization information
What are some common challenges Machine Learning Engineers face when implementing quantization techniques in production models?
What are the key skills and qualifications needed to thrive as a Machine Learning Engineer Quantization, and why are they important?
What does a Machine Learning Engineer Quantization do?
What is the difference between Machine Learning Engineer Quantization vs Data Scientist?
| Aspect | Machine Learning Engineer Quantization | Data Scientist |
|---|---|---|
| Required Credentials | Bachelor's or master's in CS, ML, or related; certifications in ML or AI | Bachelor's or master's in statistics, CS, or related; certifications in data analysis or statistics |
| Work Environment | Developing optimized ML models, deploying quantized models for efficiency | Analyzing data, building predictive models, interpreting results |
| Industry Usage | Tech companies, AI hardware firms, embedded systems | Finance, healthcare, marketing, research institutions |
Machine Learning Engineer Quantization focuses on optimizing ML models for deployment efficiency, often working closely with hardware and software teams. Data Scientists analyze data and build models for insights. While both roles require ML knowledge, quantization engineers specialize in model compression techniques, whereas data scientists focus on data analysis and interpretation.
- Volunteer Junior Machine Learning Engineer
- Machine Learning Engineer Biotech
- Machine Engineer
- Machine Learning Engineer Apprenticeship
- Manager Remote Machine Learning Engineer
- Staff Software Engineer Machine Learning
- Urgently Hiring Machine Learning Engineer New Grad
- Machine Learning Engineer
- Temporary Machine Learning R
- Senior Machine Learning Researcher
- Machine Learning Engineer Software Engineer
- Learning Disability
- Hourly Remote Machine Learning
- Machine Learning Research Engineer
- Reinforcement Learning Engineer
- Machine Learning Ai
- Freelance Machine Learning Compiler Engineer
- Artificial Intelligence Machine Learning Engineer
- Patterned Learning Ai
- Snapdragon

Full-time
Re-posted 12 days ago
Tesla rating
8.5
Based on 679 frontline employees who took The Breakroom Quiz
1st of 44 rated automakers
Job description
Tesla is a leading company in the field of AI, focusing on building foundational models for real-world autonomy. The Machine Learning Engineer will be responsible for architecting and scaling quantization pipelines for multi-modal foundation models, optimizing inference latency and power consumption for various applications including self-driving cars and robots.
Responsibilities:
• Architect and scale quantization pipelines (both Post-Training Quantization and Quantization-Aware Training) for massive multi-modal foundation models that fuse vision, prediction, and decision-making. You will optimize inference latency, memory bandwidth utilization, and power consumption for self-driving cars, Optimus robots, and digital agents operating at enterprise scale
• Innovate quantization-aware-training recipes and algorithms that tackle complex optimization challenges inherent to low-precision training
• Push the limits of low-precision AI: Research and implement advanced low-bitweight post-training quantization techniques to address hard algorithmic problems such as activation outlier mitigation, KV cache compression, and optimal layer-wise bit-allocation while strictly maintaining model accuracy
• Collaborate closely with AI compiler, inference engine, and silicon teams to ensure models are architected to maximally utilize underlying hardware capabilities by co-designing quantization-friendly architectures, hardware-aware sparsity patterns, and mixed low-precision kernels
• Collaborate across perception, planning, robotics, digital agents, and infrastructure teams to move models from research to fleet-wide, robot-wide, and enterprise-wide deployment
Qualifications:
Required:
• Degree or equivalent experience in Computer Science, Machine Learning, Robotics, Computer Vision, or related quantitative field
• 2+ years of hands-on experience training, optimizing, and deploying large-scale quantized deep learning models
• Strong technical understanding of the challenges inherent to quantizing large transformer architectures, including mitigating massive activation outliers, KV cache quantization, and maintaining the numerical stability of attention mechanisms at low precision
• Deep expertise in the theory and low-level implementation of modern quantization algorithms (e.g., GPTQ, AWQ, SmoothQuant, OmniQuant)
• Experience with low-level numerics and emerging data formats (e.g., FP8, INT4, W4A8, W8A8, micro-scaling/MX formats) and their trade-offs regarding latency, memory bandwidth, and model fidelity
• Rigorous understanding of computer architecture and the roofline model. Familiarity with how to optimize for memory hierarchies, minimize SRAM/DRAM data movement, and efficiently map quantized GEMMs and memory-bound operators to custom silicon
• Proficiency in writing custom CUDA/Triton kernels, implementing custom autograd functions (e.g., Straight-Through Estimators), and manipulating PyTorch computational graphs (e.g., FX tracing, torch.compile)
• Strong software engineering skills — clean, production-grade Python/C++ code that ships reliably at scale
• Proven ability to turn cutting-edge research into robust, real-world systems that improve safety, capability, efficiency, or digital productivity
• Passion for Tesla’s mission and excitement about deploying AI that moves both the physical and digital worlds forward
Company:
Tesla is an electric vehicle and clean energy company that provides electric cars, solar, and renewable energy solutions. Founded in 2003, the company is headquartered in Austin, USA, with a team of 10001+ employees. The company is currently Late Stage.