... CUDA, OpenMP) and experience reasoning about memory hierarchies, data movement, and compute utilization. • Demonstrated capability to collaborate directly with external partners and customers in a ...
... CUDA, OpenMP) and experience reasoning about memory hierarchies, data movement, and compute utilization. • Demonstrated capability to collaborate directly with external partners and customers in a ...
Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization
San Jose, CA · On-site
$169K - $351K/yr
Strong proficiency in C++ and Python, with solid expertise in parallel programming (CUDA, OpenMP) and low-level system profiling tools. Deep familiarity with mainstream inference engines (e.g ...
Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization
San Jose, CA · On-site
$169K - $351K/yr
Strong proficiency in C++ and Python, with solid expertise in parallel programming (CUDA, OpenMP) and low-level system profiling tools. Deep familiarity with mainstream inference engines (e.g ...
Senior Systems Software Engineer - Deep Learning Solutions
Santa Clara, CA · On-site
$65 - $83.75/hr
Background in parallel programming (e.g., CUDA, OpenMP) and experience reasoning about memory hierarchies, data movement, and compute utilization. * Demonstrated capability to collaborate directly ...
Senior Systems Software Engineer - Deep Learning Solutions
Santa Clara, CA · On-site
$65 - $83.75/hr
Background in parallel programming (e.g., CUDA, OpenMP) and experience reasoning about memory hierarchies, data movement, and compute utilization. * Demonstrated capability to collaborate directly ...
Senior Systems Software Engineer - Deep Learning Solutions
Santa Clara, CA · On-site
$143K - $189K/yr
Background in parallel programming (e.g., CUDA, OpenMP) and experience reasoning about memory hierarchies, data movement, and compute utilization. Demonstrated capability to collaborate directly with ...
Senior Systems Software Engineer - Deep Learning Solutions
Santa Clara, CA · On-site
$143K - $189K/yr
Background in parallel programming (e.g., CUDA, OpenMP) and experience reasoning about memory hierarchies, data movement, and compute utilization. Demonstrated capability to collaborate directly with ...
OpenMP, MPI, CUDA, OpenACC: these are not resume keywords, they are tools you reach for depending on the problem. At Synopsys, you will work on LS-DYNA, software that powers crash simulations ...
OpenMP, MPI, CUDA, OpenACC: these are not resume keywords, they are tools you reach for depending on the problem. At Synopsys, you will work on LS-DYNA, software that powers crash simulations ...
Senior Systems Software Engineer - Deep Learning Solutions
Santa Clara, CA · On-site
$65 - $83.75/hr
Background in parallel programming (e.g., CUDA, OpenMP) and experience reasoning about memory hierarchies, data movement, and compute utilization. * Demonstrated capability to collaborate directly ...
Senior Systems Software Engineer - Deep Learning Solutions
Santa Clara, CA · On-site
$65 - $83.75/hr
Background in parallel programming (e.g., CUDA, OpenMP) and experience reasoning about memory hierarchies, data movement, and compute utilization. * Demonstrated capability to collaborate directly ...
Coursework or project experience in parallel computing or GPU programming concepts (e.g., NVIDIA CUDA, OpenMP, or multi-threading). * Coursework or academic projects involving digital signal ...
Coursework or project experience in parallel computing or GPU programming concepts (e.g., NVIDIA CUDA, OpenMP, or multi-threading). * Coursework or academic projects involving digital signal ...
R&D Engineering, Staff Engineer - 18291
Livermore, CA · On-site
$144K/yr
OpenMP, MPI, CUDA, OpenACC: these are not resume keywords, they are tools you reach for depending on the problem. At Synopsys, you will work on LS-DYNA, software that powers crash simulations ...
R&D Engineering, Staff Engineer - 18291
Livermore, CA · On-site
$144K/yr
OpenMP, MPI, CUDA, OpenACC: these are not resume keywords, they are tools you reach for depending on the problem. At Synopsys, you will work on LS-DYNA, software that powers crash simulations ...
Senior Fortran Compiler Engineer
Santa Clara, CA · On-site
$122K - $168K/yr
... OpenMP, or CUDA • You have a real passion for compiler development Company : NVIDIA designs graphics processing units and accelerated computing hardware for AI and high-performance computing ...
Senior Fortran Compiler Engineer
Santa Clara, CA · On-site
$122K - $168K/yr
... OpenMP, or CUDA • You have a real passion for compiler development Company : NVIDIA designs graphics processing units and accelerated computing hardware for AI and high-performance computing ...
Background in parallel programming, e.g., CUDA, OpenMP, MPI, pthreads, etc. * Programming fluency in C/C++ with a deep understanding of algorithms and software development. * Knowledge of CPU and GPU ...
Background in parallel programming, e.g., CUDA, OpenMP, MPI, pthreads, etc. * Programming fluency in C/C++ with a deep understanding of algorithms and software development. * Knowledge of CPU and GPU ...
Background in parallel programming, e.g., CUDA, OpenMP, MPI, pthreads, etc. * Programming fluency in C/C++ with a deep understanding of algorithms and software development. * Knowledge of CPU and GPU ...
Background in parallel programming, e.g., CUDA, OpenMP, MPI, pthreads, etc. * Programming fluency in C/C++ with a deep understanding of algorithms and software development. * Knowledge of CPU and GPU ...
Software Engineer
Lexington, MA · On-site
OpenMP, Threads, Real-Time Priorities, CUDA, FPGA * Python/TensorFlow * Web Application Development Software * Cloud-Based Solutions (e.g. AWS, Azure, Google) Background/Need The Secure Resilient ...
Software Engineer
Lexington, MA · On-site
OpenMP, Threads, Real-Time Priorities, CUDA, FPGA * Python/TensorFlow * Web Application Development Software * Cloud-Based Solutions (e.g. AWS, Azure, Google) Background/Need The Secure Resilient ...
In-depth expertise in linear algebra and performance optimization of Deep Learning training and inference Background in parallel programming, e.g., CUDA, OpenMP, MPI, pthreads, etc Programming ...
In-depth expertise in linear algebra and performance optimization of Deep Learning training and inference Background in parallel programming, e.g., CUDA, OpenMP, MPI, pthreads, etc Programming ...
In-depth expertise in linear algebra and performance optimization of Deep Learning training and inference Background in parallel programming, e.g., CUDA, OpenMP, MPI, pthreads, etc Programming ...
In-depth expertise in linear algebra and performance optimization of Deep Learning training and inference Background in parallel programming, e.g., CUDA, OpenMP, MPI, pthreads, etc Programming ...
Background in parallel programming, e.g., CUDA, OpenMP, MPI, pthreads, etc. * Programming fluency in C/C++ with a deep understanding of algorithms and software development. * Knowledge of CPU and GPU ...
Background in parallel programming, e.g., CUDA, OpenMP, MPI, pthreads, etc. * Programming fluency in C/C++ with a deep understanding of algorithms and software development. * Knowledge of CPU and GPU ...
Background in parallel programming, e.g., CUDA, OpenMP, MPI, pthreads, etc. * Programming fluency in C/C++ with a deep understanding of algorithms and software development. * Knowledge of CPU and GPU ...
Background in parallel programming, e.g., CUDA, OpenMP, MPI, pthreads, etc. * Programming fluency in C/C++ with a deep understanding of algorithms and software development. * Knowledge of CPU and GPU ...
In-depth expertise in linear algebra and performance optimization of Deep Learning training and inference Background in parallel programming, e.g., CUDA, OpenMP, MPI, pthreads, etc Programming ...
In-depth expertise in linear algebra and performance optimization of Deep Learning training and inference Background in parallel programming, e.g., CUDA, OpenMP, MPI, pthreads, etc Programming ...
In-depth expertise in linear algebra and performance optimization of Deep Learning training and inference Background in parallel programming, e.g., CUDA, OpenMP, MPI, pthreads, etc Programming ...
In-depth expertise in linear algebra and performance optimization of Deep Learning training and inference Background in parallel programming, e.g., CUDA, OpenMP, MPI, pthreads, etc Programming ...
In-depth expertise in linear algebra and performance optimization of Deep Learning training and inference Background in parallel programming, e.g., CUDA, OpenMP, MPI, pthreads, etc Programming ...
In-depth expertise in linear algebra and performance optimization of Deep Learning training and inference Background in parallel programming, e.g., CUDA, OpenMP, MPI, pthreads, etc Programming ...
Background in parallel programming, e.g., CUDA, OpenMP, MPI, pthreads, etc. * Programming fluency in C/C++ with a deep understanding of algorithms and software development. * Knowledge of CPU and GPU ...
Background in parallel programming, e.g., CUDA, OpenMP, MPI, pthreads, etc. * Programming fluency in C/C++ with a deep understanding of algorithms and software development. * Knowledge of CPU and GPU ...
Cuda Openmp information
See salary details
$111.5K - $120.1K
0% of jobs
$120.1K - $128.7K
0% of jobs
$128.7K - $137.3K
0% of jobs
$137.3K - $145.9K
4% of jobs
$145.9K - $154.5K
0% of jobs
$154.5K - $163K
0% of jobs
$163K - $171.6K
0% of jobs
$171.6K - $180.2K
2% of jobs
$180.2K - $188.8K
0% of jobs
$188.8K - $197.4K
0% of jobs
$199.1K is the 25th percentile. Wages below this are outliers.
$197.4K - $206K
94% of jobs
$111.5K
$206K
How much do cuda openmp jobs pay per year?
What is the difference between Cuda Openmp vs C++ Developer?
| Aspect | Cuda Openmp | C++ Developer |
|---|---|---|
| Required Credentials | Knowledge of GPU programming, CUDA, OpenMP, parallel computing | Proficiency in C++, algorithms, software development |
| Work Environment | High-performance computing, parallel processing environments | Software development across various industries |
| Industry Usage | Research, scientific computing, AI, HPC | Software engineering, application development, system programming |
While Cuda Openmp focuses on parallel programming for GPU and CPU acceleration, C++ Developers create software applications across multiple domains. Both roles require strong programming skills, but Cuda Openmp specialists emphasize parallel computing techniques, whereas C++ Developers focus on software design and implementation.
What other helpful pages are available for Cuda Openmp?
Other pages related to Cuda Openmp:

Senior Systems Software Engineer - Deep Learning Solutions
Santa Clara, CA • On-site
Full-time
Re-posted 6 days ago
Nvidia rating
9.6
Based on 18 frontline employees who took The Breakroom Quiz
Job description
NVIDIA is a global leader in physical AI, powering self-driving cars, humanoid robots, intelligent environments, and medical devices. We are hiring a Senior Systems Software Engineer to optimize deep learning inference for autonomous vehicles and robotics on edge devices, collaborating with various teams and external partners to enhance performance and deliver solutions.
Responsibilities:
• Address customer and partner optimization challenges: Engage directly with prominent automotive OEMs and robotics associates to analyze, debug, and improve their deep learning models on NVIDIA platforms. We emphasize delivering solutions rather than just recommendations.
• Own performance benchmarking: Drive efforts to achieve leading results on MLPerf Edge and industry benchmarks, as well as closed-source engagements with key partners. Define methodology, ensure reproducibility, and turn results into actionable optimization priorities.
• Evaluate emerging model architectures: Investigate new DL architectures, including vision encoders, multi-modal VLMs, hybrid SSM-Transformer backbones, diffusion/flow matching decoders, and multi-camera tokenizers, regarding compilation feasibility, memory footprint, and latency on target SOCs.
• Collaborate across teams: Work alongside our compiler, runtime, and hardware groups to link model-level insight with platform capabilities.
• Contribute to build reviews and help develop internal roadmap priorities based on real customer workload patterns.
• Represent NVIDIA externally: Share our deep learning optimization expertise at conferences, webinars, and partner events. Help elevate the broader team by bringing back insights and establishing guidelines.
• Deliver TensorRT and compiler-stack solutions for edge: Build and deploy inference solutions on Jetson, DRIVE, and GPU + ARM platforms for AV and robotics workloads. Develop Proofs of Readiness (PORs) and collaborate closely with our compiler team on Torch-TRT, MLIR-TRT, and related frameworks to bridge performance gaps.
Qualifications:
Required:
• Master’s degree or equivalent experience in Computer Science, Electrical Engineering, or a related field.
• Over 12 years working in the industry, including at least 8 years specializing in deep learning model optimization, inference engineering, or neural network compilation. Proficiency in understanding and reviewing model architectures at the operator/kernel level, not merely handling their operation, is required.
• Over 5 years of validated expertise in embedded/edge software, with experience delivering production inference solutions within power-limited, latency-sensitive deployment environments.
• Comprehensive knowledge of contemporary DL architectures: transformers, attention variants, vision encoders (ViT), multi-modal/vision-language model frameworks, as well as experience with diffusion models and/or state space models.
• Expert knowledge of GPU architecture fundamentals, CUDA, and low-level performance optimization using heterogeneous computing. Experience with TensorRT, compiler IRs, or equivalent inference optimization toolchains.
• Solid understanding of embedded operating system internals (QNX/Linux), memory management, C/C++, and embedded/system software concepts.
• Background in parallel programming (e.g., CUDA, OpenMP) and experience reasoning about memory hierarchies, data movement, and compute utilization.
• Demonstrated capability to collaborate directly with external partners and customers in a deep technical role. You solve their workload issues, identify performance problems, and provide solutions within production limitations.
Preferred:
• Experience with ML compiler frameworks (TVM, MLIR, XLA, Triton) or contributing to inference runtime development.
• Production deployment experience with autonomous vehicle perception or planning stacks, understanding the full pipeline from sensor input through trajectory output.
• Familiarity with the Physical AI model landscape: VLM + action expert architectures, end-to-end driving models, or robot foundation models.
• Contributions to MLPerf benchmarks and large-scale industry performance optimization efforts.
• Experience with automotive safety standards (ISO 26262, SOTIF) and their implications for inference system development.
Company:
NVIDIA designs graphics processing units and accelerated computing hardware for AI and high-performance computing applications. Founded in 1993, the company is headquartered in Santa Clara, USA, with a team of 10001+ employees. The company is currently Late Stage.
About Nvidia
Sourced by ZipRecruiter
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology--and amazing people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what's never been done before takes vision, innovation, and the world's best talent.
Industry
Computer and electronic product manufacturing
Company size
10,000+ Employees
Headquarters location
Santa Clara, CA, US