1

Pytorch Onnx Jobs (NOW HIRING)

Build and maintain model ingestion pipelines (e.g., PyTorch / ONNX → internal IR) * Implement graph transformations such as: * Operator decomposition and canonicalization * Shape inference and ...

Kubernetes, Ray, vLLM and PyTorch, onnx etc. * Strong programming skills in Python, C/C++, CUDA. * Excellent communication skills.

Build and maintain model ingestion pipelines (e.g., PyTorch / ONNX → internal IR) * Implement graph transformations such as: * Operator decomposition and canonicalization * Shape inference and ...

Guide model porting across frameworks and runtimes such as PyTorch, ONNX, and vendor-specific runtimes * Build prototypes and proofs of concept to reduce technical risk before engineering investment

New

... PyTorch, ONNX, and torch.compile, including model rewrites and graph-level transformations. • Design and implement fusion kernels using DSL based approaches (e.g., Triton), enabling fused ...

Experience in developing or using deep learning frameworks (e.g., TensorFlow, Kera's, PyTorch, ONNX, etc.) * Experience in optimization of deep learning/ML algorithms (e.g., Retraining, Quantization ...

... PyTorch, ONNX) and graph-level optimizations • Familiarity with ML hardware accelerators, memory hierarchies, and performance modeling • Solid C++ programming skills • Proven ability to lead ...

Experience deploying or optimizing neural network inference workloads using technologies such as PyTorch, ONNX, and TensorRT . * Experience with real-time embedded systems and handling large data ...

Experience deploying or optimizing neural network inference workloads using technologies such as PyTorch, ONNX, and TensorRT . * Experience with real-time embedded systems and handling large data ...

... PyTorch ONNX vendor specific runtimes) Build prototypes and proof of concepts to reduce technical risk prior to full engineering investment Qualifications Required Bachelor's or Master's degree in ...

Drive graph capture and deployment using PyTorch, ONNX, and torch.compile, including model rewrites and graph-level transformations. * Design and implement fusion kernels using DSL based approaches ...

next page

Showing results 1-20

Pytorch Onnx information

What is PyTorch ONNX?

PyTorch ONNX refers to the process of exporting models built in PyTorch to the Open Neural Network Exchange (ONNX) format. This allows models trained in PyTorch to be used in different frameworks and environments that support ONNX, such as TensorFlow, Caffe2, or various deployment tools. The ONNX format provides interoperability and flexibility, making it easier to deploy machine learning models across different platforms. PyTorch provides built-in functions like `torch.onnx.export()` to facilitate this conversion. Using ONNX can help streamline workflows for developers working in production or research settings.

What are common challenges when converting models from PyTorch to ONNX?

When converting PyTorch models to ONNX, professionals often encounter challenges like unsupported operators, dynamic input shapes, or differences in layer implementations between frameworks. These issues may require modifying the original PyTorch model or using custom export functions to ensure compatibility. Collaboration between data scientists and software engineers is crucial to troubleshoot export errors and validate model performance in the target ONNX runtime environment, ensuring the model behaves consistently across platforms.

What skills and qualifications are needed to work with PyTorch ONNX?

To thrive as a PyTorch ONNX Engineer, you need a strong background in deep learning, proficiency in Python programming, and experience with PyTorch and model optimization techniques. Familiarity with ONNX (Open Neural Network Exchange), model conversion workflows, and deployment tools like TensorRT or ONNX Runtime is typically required. Analytical thinking, problem-solving, and clear communication are key soft skills that help in collaborating with cross-functional teams and troubleshooting complex issues. These skills are essential for successfully building, optimizing, and deploying machine learning models across diverse platforms.

What is the difference between Pytorch Onnx vs Machine Learning Engineer?

AspectPytorch OnnxMachine Learning Engineer
Primary RoleModel conversion and interoperabilityDeveloping, deploying, and optimizing ML models
Skills RequiredDeep learning frameworks, model export, PythonML algorithms, programming, data analysis
Work EnvironmentAI/ML teams, software developmentData science teams, product development
Industry UsageModel deployment, cross-platform compatibilityModel development, research, and deployment

While Pytorch Onnx focuses on converting and deploying models across platforms, Machine Learning Engineers design and optimize models for various applications. Both roles require knowledge of ML frameworks and programming, but their core responsibilities differ significantly.

AI Performance Engineer (Cloud AI Engineering), Sr | Staff | Sr. Staff

San Diego, CA • On-site

Qualcomm
Technology, Communication and Media • 10K+ employees

Full-time

Re-posted 3 days ago


Qualcomm rating

8.8

Company rating: 8.8 out of 10

Based on 11 frontline employees who took The Breakroom Quiz

48th of 247 rated software companies


Job description

Job Summary:
Qualcomm Technologies, Inc. is at the forefront of Cloud AI, leveraging its expertise in digital wireless technologies. They are seeking an AI Performance Engineer to optimize and deploy models for efficient inference, working closely with customers and internal teams to enhance AI workloads.
Responsibilities:
• Convert, optimize and deploy models for efficient inference using PyTorch, ONNX.
• Work at the forefront of GenAI by understanding advanced algorithms (e.g. attention mechanisms, MoEs) and numerics to identify new optimization opportunities.
• Performance analysis and optimization of LLM, VLM, and diffusion models for inference. Scale performance for throughput and latency constraints.
• Mapping the next generation AI workloads on top of current and future hardware designs.
• Work closely with customers to drive solutions by collaborating with internal compiler, firmware and platform teams.
• Analyze complex performance or stability issues to work towards final root cause of underlying problems.
• Create engineering solutions to deliver continuous insights into performance of AI workloads guiding the improvements over time.
• Design and implement high-level kernels, e.g. in Triton, with a focus on generating efficient, low-level code.
Qualifications:
Required:
• Hands-on experience in building and optimizing language models, notably in PyTorch, ONNX, preferably in production-grade environments.
• Deep understanding of transformer architectures, attention mechanisms and performance trade-offs.
• Experience in workload mapping strategies exhibiting sharding or various parallelisms.
• Strong Python programming skills.
• Proactive learning about the latest inference optimization techniques.
• Understanding of computer architecture, ML accelerators, in-memory processing and distributed systems.
• Strong communication, problem-solving skills and ability to learn and work effectively in a fast-paced and collaborative environment.
• MS in Computer Science, Machine Learning, Computer Engineering or Electrical Engineering.
• Bachelor's degree in Computer Science, Engineering, Information Systems, or related field and 6+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.
• Master's degree in Computer Science, Engineering, Information Systems, or related field and 5+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.
• PhD in Computer Science, Engineering, Information Systems, or related field and 4+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.
Preferred:
• Background in neural network operators and mathematical operations, including linear algebra and math libraries.
• Understanding of machine learning compilers.
• Experience in converging accuracy and its evaluation methods.
• Knowledge of torch.compile or torchDynamo.
• PhD in Computer Science, Computer Engineering or Machine Learning
Company:
Qualcomm designs wireless technologies and semiconductors that power connectivity, communication, and smart devices. Founded in 1985, the company is headquartered in San Diego, USA, with a team of 10001+ employees. The company is currently Late Stage.

What Qualcomm employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom


Qualcomm logo

About Qualcomm

Sourced by ZipRecruiter

Qualcomm is enabling a world where everyone and everything can be intelligently connected. You interact with products and technologies made possible by Qualcomm every day, including 5G-enabled smartphones that double as pro-level cameras and gaming devices, smarter vehicles and cities, and the technology behind the smart, connected factories that manufactured your latest purchase. Our powerful connectivity solutions keep you connected—even in remote areas. Qualcomm 5G and AI innovations are the power behind the connected intelligent edge. You’ll find our technologies behind and inside the innovations that deliver significant value across multiple industries and to billions of people every day.

Industry

Technology, communication and media

Company size

10,000+ Employees

Headquarters location

San Diego, CA, US

Year founded

1985