Responsibilities : • Convert, optimize and deploy models for efficient inference using PyTorch, ONNX. • Work at the forefront of GenAI by understanding advanced algorithms (e.g. attention ...
Responsibilities : • Convert, optimize and deploy models for efficient inference using PyTorch, ONNX. • Work at the forefront of GenAI by understanding advanced algorithms (e.g. attention ...
ML Compiler Engineer
San Bruno, CA · On-site
Build and maintain model ingestion pipelines (e.g., PyTorch / ONNX → internal IR) * Implement graph transformations such as: * Operator decomposition and canonicalization * Shape inference and ...
ML Compiler Engineer
San Bruno, CA · On-site
Build and maintain model ingestion pipelines (e.g., PyTorch / ONNX → internal IR) * Implement graph transformations such as: * Operator decomposition and canonicalization * Shape inference and ...
Kubernetes, Ray, vLLM and PyTorch, onnx etc. * Strong programming skills in Python, C/C++, CUDA. * Excellent communication skills.
Kubernetes, Ray, vLLM and PyTorch, onnx etc. * Strong programming skills in Python, C/C++, CUDA. * Excellent communication skills.
ML Compiler Engineer
San Bruno, CA · On-site
Build and maintain model ingestion pipelines (e.g., PyTorch / ONNX → internal IR) * Implement graph transformations such as: * Operator decomposition and canonicalization * Shape inference and ...
ML Compiler Engineer
San Bruno, CA · On-site
Build and maintain model ingestion pipelines (e.g., PyTorch / ONNX → internal IR) * Implement graph transformations such as: * Operator decomposition and canonicalization * Shape inference and ...
... PyTorch, ONNX, JAX Preferred : • 1+ years Python programming experience • Experience with different NN architectures: DNNs, CNNs, RNNs/LSTMs, GANs, LLMs, etc. • Experience with Graph ...
... PyTorch, ONNX, JAX Preferred : • 1+ years Python programming experience • Experience with different NN architectures: DNNs, CNNs, RNNs/LSTMs, GANs, LLMs, etc. • Experience with Graph ...
AI Architect - Media
Atlanta, GA · On-site
Guide model porting across frameworks and runtimes such as PyTorch, ONNX, and vendor-specific runtimes * Build prototypes and proofs of concept to reduce technical risk before engineering investment
New
AI Architect - Media
Atlanta, GA · On-site
Guide model porting across frameworks and runtimes such as PyTorch, ONNX, and vendor-specific runtimes * Build prototypes and proofs of concept to reduce technical risk before engineering investment
New
AI Performance Engineer (Cloud AI Engineering), Sr | Staff | Sr. Staff
San Diego, CA · On-site
$110K - $152K/yr
Convert, optimize and deploy models for efficient inference using PyTorch, ONNX. * Work at the forefront of GenAI by understanding advanced algorithms (e.g. attention mechanisms, MoEs) and numerics ...
AI Performance Engineer (Cloud AI Engineering), Sr | Staff | Sr. Staff
San Diego, CA · On-site
$110K - $152K/yr
Convert, optimize and deploy models for efficient inference using PyTorch, ONNX. * Work at the forefront of GenAI by understanding advanced algorithms (e.g. attention mechanisms, MoEs) and numerics ...
Senior Software Engineer - Prediction
Pittsburgh, PA · On-site
$118K - $156K/yr
... PyTorch, ONNX and deployment/optimization of models is desired • Industry experience writing production-quality, performance-critical code, and maintaining large codebases is desired Company
Senior Software Engineer - Prediction
Pittsburgh, PA · On-site
$118K - $156K/yr
... PyTorch, ONNX and deployment/optimization of models is desired • Industry experience writing production-quality, performance-critical code, and maintaining large codebases is desired Company
PyTorch, ONNX/TensorRT export, and the tooling that keeps annotation, training, and deployment moving.* Partner with the edge team to quantize and prune models for on-device inference.## What you'll ...
PyTorch, ONNX/TensorRT export, and the tooling that keeps annotation, training, and deployment moving.* Partner with the edge team to quantize and prune models for on-device inference.## What you'll ...
... PyTorch, ONNX, and torch.compile, including model rewrites and graph-level transformations. • Design and implement fusion kernels using DSL based approaches (e.g., Triton), enabling fused ...
... PyTorch, ONNX, and torch.compile, including model rewrites and graph-level transformations. • Design and implement fusion kernels using DSL based approaches (e.g., Triton), enabling fused ...
Experience working with one of the Deep Learning frameworks like TensorFlow, PyTorch, ONNX, JAX Preferred Qualifications * 1+ years Python programming experience * Experience with different NN ...
Experience working with one of the Deep Learning frameworks like TensorFlow, PyTorch, ONNX, JAX Preferred Qualifications * 1+ years Python programming experience * Experience with different NN ...
Experience in developing or using deep learning frameworks (e.g., TensorFlow, Kera's, PyTorch, ONNX, etc.) * Experience in optimization of deep learning/ML algorithms (e.g., Retraining, Quantization ...
Experience in developing or using deep learning frameworks (e.g., TensorFlow, Kera's, PyTorch, ONNX, etc.) * Experience in optimization of deep learning/ML algorithms (e.g., Retraining, Quantization ...
Machine Learning Compiler
Raleigh, NC · On-site
... PyTorch, ONNX) and graph-level optimizations • Familiarity with ML hardware accelerators, memory hierarchies, and performance modeling • Solid C++ programming skills • Proven ability to lead ...
Machine Learning Compiler
Raleigh, NC · On-site
... PyTorch, ONNX) and graph-level optimizations • Familiarity with ML hardware accelerators, memory hierarchies, and performance modeling • Solid C++ programming skills • Proven ability to lead ...
GPU Engineer
Houston, TX · On-site
Experience deploying or optimizing neural network inference workloads using technologies such as PyTorch, ONNX, and TensorRT . * Experience with real-time embedded systems and handling large data ...
GPU Engineer
Houston, TX · On-site
Experience deploying or optimizing neural network inference workloads using technologies such as PyTorch, ONNX, and TensorRT . * Experience with real-time embedded systems and handling large data ...
AI Architect - Media
Atlanta, GA · On-site
Guide model porting across frameworks and runtimes (e.g., PyTorch ONNX vendor specific runtimes) * Build prototypes and proof of concepts to reduce technical risk prior to full engineering investment ...
AI Architect - Media
Atlanta, GA · On-site
Guide model porting across frameworks and runtimes (e.g., PyTorch ONNX vendor specific runtimes) * Build prototypes and proof of concepts to reduce technical risk prior to full engineering investment ...
GPU Engineer
Houston, TX · On-site
Experience deploying or optimizing neural network inference workloads using technologies such as PyTorch, ONNX, and TensorRT . * Experience with real-time embedded systems and handling large data ...
GPU Engineer
Houston, TX · On-site
Experience deploying or optimizing neural network inference workloads using technologies such as PyTorch, ONNX, and TensorRT . * Experience with real-time embedded systems and handling large data ...
AI Architect - Media
Atlanta, GA · On-site
... PyTorch ONNX vendor specific runtimes) Build prototypes and proof of concepts to reduce technical risk prior to full engineering investment Qualifications Required Bachelor's or Master's degree in ...
AI Architect - Media
Atlanta, GA · On-site
... PyTorch ONNX vendor specific runtimes) Build prototypes and proof of concepts to reduce technical risk prior to full engineering investment Qualifications Required Bachelor's or Master's degree in ...
AI Model Optimization Architect
San Diego, CA · On-site
$158K - $237K/yr
Drive graph capture and deployment using PyTorch, ONNX, and torch.compile, including model rewrites and graph-level transformations. * Design and implement fusion kernels using DSL based approaches ...
AI Model Optimization Architect
San Diego, CA · On-site
$158K - $237K/yr
Drive graph capture and deployment using PyTorch, ONNX, and torch.compile, including model rewrites and graph-level transformations. * Design and implement fusion kernels using DSL based approaches ...
Experience in ML frameworks (TensorFlow, PyTorch, ONNX) * Experience in micro architecture of SoC, interface subsystems, logic modules * Self-motivated problem-solver with an ability to work well in ...
Experience in ML frameworks (TensorFlow, PyTorch, ONNX) * Experience in micro architecture of SoC, interface subsystems, logic modules * Self-motivated problem-solver with an ability to work well in ...
Prior experience with AI frameworks and engines, such as TensorRT, PyTorch, ONNX, OpenVINO, vLLM, or TRT-LLM. * Knowledge of GPU memory management, cache management, or high-performance networking.
Prior experience with AI frameworks and engines, such as TensorRT, PyTorch, ONNX, OpenVINO, vLLM, or TRT-LLM. * Knowledge of GPU memory management, cache management, or high-performance networking.
Pytorch Onnx information
What is PyTorch ONNX?
What are common challenges when converting models from PyTorch to ONNX?
What skills and qualifications are needed to work with PyTorch ONNX?
What is the difference between Pytorch Onnx vs Machine Learning Engineer?
| Aspect | Pytorch Onnx | Machine Learning Engineer |
|---|---|---|
| Primary Role | Model conversion and interoperability | Developing, deploying, and optimizing ML models |
| Skills Required | Deep learning frameworks, model export, Python | ML algorithms, programming, data analysis |
| Work Environment | AI/ML teams, software development | Data science teams, product development |
| Industry Usage | Model deployment, cross-platform compatibility | Model development, research, and deployment |
While Pytorch Onnx focuses on converting and deploying models across platforms, Machine Learning Engineers design and optimize models for various applications. Both roles require knowledge of ML frameworks and programming, but their core responsibilities differ significantly.
AI Performance Engineer (Cloud AI Engineering), Sr | Staff | Sr. Staff
San Diego, CA • On-site
Full-time
Re-posted 3 days ago
Qualcomm rating
8.8
Based on 11 frontline employees who took The Breakroom Quiz
48th of 247 rated software companies
Job description
Qualcomm Technologies, Inc. is at the forefront of Cloud AI, leveraging its expertise in digital wireless technologies. They are seeking an AI Performance Engineer to optimize and deploy models for efficient inference, working closely with customers and internal teams to enhance AI workloads.
Responsibilities:
• Convert, optimize and deploy models for efficient inference using PyTorch, ONNX.
• Work at the forefront of GenAI by understanding advanced algorithms (e.g. attention mechanisms, MoEs) and numerics to identify new optimization opportunities.
• Performance analysis and optimization of LLM, VLM, and diffusion models for inference. Scale performance for throughput and latency constraints.
• Mapping the next generation AI workloads on top of current and future hardware designs.
• Work closely with customers to drive solutions by collaborating with internal compiler, firmware and platform teams.
• Analyze complex performance or stability issues to work towards final root cause of underlying problems.
• Create engineering solutions to deliver continuous insights into performance of AI workloads guiding the improvements over time.
• Design and implement high-level kernels, e.g. in Triton, with a focus on generating efficient, low-level code.
Qualifications:
Required:
• Hands-on experience in building and optimizing language models, notably in PyTorch, ONNX, preferably in production-grade environments.
• Deep understanding of transformer architectures, attention mechanisms and performance trade-offs.
• Experience in workload mapping strategies exhibiting sharding or various parallelisms.
• Strong Python programming skills.
• Proactive learning about the latest inference optimization techniques.
• Understanding of computer architecture, ML accelerators, in-memory processing and distributed systems.
• Strong communication, problem-solving skills and ability to learn and work effectively in a fast-paced and collaborative environment.
• MS in Computer Science, Machine Learning, Computer Engineering or Electrical Engineering.
• Bachelor's degree in Computer Science, Engineering, Information Systems, or related field and 6+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.
• Master's degree in Computer Science, Engineering, Information Systems, or related field and 5+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.
• PhD in Computer Science, Engineering, Information Systems, or related field and 4+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.
Preferred:
• Background in neural network operators and mathematical operations, including linear algebra and math libraries.
• Understanding of machine learning compilers.
• Experience in converging accuracy and its evaluation methods.
• Knowledge of torch.compile or torchDynamo.
• PhD in Computer Science, Computer Engineering or Machine Learning
Company:
Qualcomm designs wireless technologies and semiconductors that power connectivity, communication, and smart devices. Founded in 1985, the company is headquartered in San Diego, USA, with a team of 10001+ employees. The company is currently Late Stage.
About Qualcomm
Sourced by ZipRecruiter
Qualcomm is enabling a world where everyone and everything can be intelligently connected. You interact with products and technologies made possible by Qualcomm every day, including 5G-enabled smartphones that double as pro-level cameras and gaming devices, smarter vehicles and cities, and the technology behind the smart, connected factories that manufactured your latest purchase. Our powerful connectivity solutions keep you connected—even in remote areas. Qualcomm 5G and AI innovations are the power behind the connected intelligent edge. You’ll find our technologies behind and inside the innovations that deliver significant value across multiple industries and to billions of people every day.
Industry
Technology, communication and media
Company size
10,000+ Employees
Headquarters location
San Diego, CA, US
Year founded
1985