1

Cuda Openmp Jobs (NOW HIRING)

Senior Fortran Compiler Engineer

Santa Clara, CA · On-site

$122K - $168K/yr

... OpenMP, or CUDA • You have a real passion for compiler development Company : NVIDIA designs graphics processing units and accelerated computing hardware for AI and high-performance computing ...

OpenMP, Threads, Real-Time Priorities, CUDA, FPGA * Python/TensorFlow * Web Application Development Software * Cloud-Based Solutions (e.g. AWS, Azure, Google) Background/Need The Secure Resilient ...

Showing results 21-40

Cuda Openmp information

See salary details

$111.5K

$206K

How much do cuda openmp jobs pay per year?

As of Sep 14, 2026, the average yearly pay for cuda openmp in the United States is $200,510.00, according to ZipRecruiter salary data. Most workers in this role earn between $205,000.00 and $205,000.00 per year, depending on experience, location, and employer.

What is the difference between Cuda Openmp vs C++ Developer?

AspectCuda OpenmpC++ Developer
Required CredentialsKnowledge of GPU programming, CUDA, OpenMP, parallel computingProficiency in C++, algorithms, software development
Work EnvironmentHigh-performance computing, parallel processing environmentsSoftware development across various industries
Industry UsageResearch, scientific computing, AI, HPCSoftware engineering, application development, system programming

While Cuda Openmp focuses on parallel programming for GPU and CPU acceleration, C++ Developers create software applications across multiple domains. Both roles require strong programming skills, but Cuda Openmp specialists emphasize parallel computing techniques, whereas C++ Developers focus on software design and implementation.

What other helpful pages are available for Cuda Openmp?

Other pages related to Cuda Openmp:

Infographic showing various Cuda Openmp job openings in the United States as of September 2026, with employment types broken down into 98% Full Time, 1% Part Time, and 1% Contract. Highlights an 77% Physical, 5% Hybrid, and 18% Remote job distribution, with an average salary of $200,510 per year, or $96.4 per hour.

Senior Systems Software Engineer - Deep Learning Solutions

Santa Clara, CA • On-site

NVIDIA
Computer and Electronic Product Manufacturing • 10K+ employees

Full-time

Re-posted 6 days ago


Nvidia rating

9.6

Company rating: 9.6 out of 10

Based on 18 frontline employees who took The Breakroom Quiz


Job description

Job Summary:
NVIDIA is a global leader in physical AI, powering self-driving cars, humanoid robots, intelligent environments, and medical devices. We are hiring a Senior Systems Software Engineer to optimize deep learning inference for autonomous vehicles and robotics on edge devices, collaborating with various teams and external partners to enhance performance and deliver solutions.
Responsibilities:
• Address customer and partner optimization challenges: Engage directly with prominent automotive OEMs and robotics associates to analyze, debug, and improve their deep learning models on NVIDIA platforms. We emphasize delivering solutions rather than just recommendations.
• Own performance benchmarking: Drive efforts to achieve leading results on MLPerf Edge and industry benchmarks, as well as closed-source engagements with key partners. Define methodology, ensure reproducibility, and turn results into actionable optimization priorities.
• Evaluate emerging model architectures: Investigate new DL architectures, including vision encoders, multi-modal VLMs, hybrid SSM-Transformer backbones, diffusion/flow matching decoders, and multi-camera tokenizers, regarding compilation feasibility, memory footprint, and latency on target SOCs.
• Collaborate across teams: Work alongside our compiler, runtime, and hardware groups to link model-level insight with platform capabilities.
• Contribute to build reviews and help develop internal roadmap priorities based on real customer workload patterns.
• Represent NVIDIA externally: Share our deep learning optimization expertise at conferences, webinars, and partner events. Help elevate the broader team by bringing back insights and establishing guidelines.
• Deliver TensorRT and compiler-stack solutions for edge: Build and deploy inference solutions on Jetson, DRIVE, and GPU + ARM platforms for AV and robotics workloads. Develop Proofs of Readiness (PORs) and collaborate closely with our compiler team on Torch-TRT, MLIR-TRT, and related frameworks to bridge performance gaps.
Qualifications:
Required:
• Master’s degree or equivalent experience in Computer Science, Electrical Engineering, or a related field.
• Over 12 years working in the industry, including at least 8 years specializing in deep learning model optimization, inference engineering, or neural network compilation. Proficiency in understanding and reviewing model architectures at the operator/kernel level, not merely handling their operation, is required.
• Over 5 years of validated expertise in embedded/edge software, with experience delivering production inference solutions within power-limited, latency-sensitive deployment environments.
• Comprehensive knowledge of contemporary DL architectures: transformers, attention variants, vision encoders (ViT), multi-modal/vision-language model frameworks, as well as experience with diffusion models and/or state space models.
• Expert knowledge of GPU architecture fundamentals, CUDA, and low-level performance optimization using heterogeneous computing. Experience with TensorRT, compiler IRs, or equivalent inference optimization toolchains.
• Solid understanding of embedded operating system internals (QNX/Linux), memory management, C/C++, and embedded/system software concepts.
• Background in parallel programming (e.g., CUDA, OpenMP) and experience reasoning about memory hierarchies, data movement, and compute utilization.
• Demonstrated capability to collaborate directly with external partners and customers in a deep technical role. You solve their workload issues, identify performance problems, and provide solutions within production limitations.
Preferred:
• Experience with ML compiler frameworks (TVM, MLIR, XLA, Triton) or contributing to inference runtime development.
• Production deployment experience with autonomous vehicle perception or planning stacks, understanding the full pipeline from sensor input through trajectory output.
• Familiarity with the Physical AI model landscape: VLM + action expert architectures, end-to-end driving models, or robot foundation models.
• Contributions to MLPerf benchmarks and large-scale industry performance optimization efforts.
• Experience with automotive safety standards (ISO 26262, SOTIF) and their implications for inference system development.
Company:
NVIDIA designs graphics processing units and accelerated computing hardware for AI and high-performance computing applications. Founded in 1993, the company is headquartered in Santa Clara, USA, with a team of 10001+ employees. The company is currently Late Stage.

What Nvidia employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom


Nvidia logo

About Nvidia

Sourced by ZipRecruiter

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology--and amazing people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what's never been done before takes vision, innovation, and the world's best talent.

Industry

Computer and electronic product manufacturing

Company size

10,000+ Employees

Headquarters location

Santa Clara, CA, US