1

Cuda Kernel Engineer Jobs in West Roxbury, MA (NOW HIRING)

Machine Learning Systems Engineer

Boston, MA · On-site +1

$144K - $192K/yr

Kernel Development : Design and maintain high-performance GPU kernels in Triton or CUDA for state-of-the-art ML workloads. * Data Pipeline Engineering : Optimize robust data loading pipelines that ...

Machine Learning Systems Engineer

Boston, MA · On-site +1

$144K - $192K/yr

Kernel Development : Design and maintain high-performance GPU kernels in Triton or CUDA for state-of-the-art ML workloads. * Data Pipeline Engineering : Optimize robust data loading pipelines that ...

Experience with CUDA programming; Experience programming distributed systems; Experience with ... Advanced knowledge of Linux kernel internals and systems programming methods; Advanced knowledge of ...

Systems Software Engineer

Boston, MA · On-site

$120K - $170K/yr

... kernel schedulers. * 3+ years of startup experience - you know the trade-offs between shipping fast ... CUDA and/or Jetson platform * Experience with classical computer vision techniques and machine ...

Systems Software Engineer

Boston, MA · On-site

$120K - $170K/yr

... kernel schedulers. * 3+ years of startup experience - you know the trade-offs between shipping fast ... CUDA and/or Jetson platform * Experience with classical computer vision techniques and machine ...

... kernel schedulers. * 3+ years of startup experience - you know the trade-offs between shipping fast ... CUDA and/or Jetson platform * Experience with classical computer vision techniques and machine ...

Showing results 21-35

Cuda Kernel Engineer information

What are common challenges faced by CUDA Kernel Engineers when optimizing GPU code for performance?

Cuda Kernel Engineers often encounter challenges such as managing memory hierarchy efficiently, minimizing data transfer between host and device, and avoiding thread divergence. Ensuring optimal occupancy and maximizing parallelism while preventing bottlenecks like bank conflicts or uncoalesced memory access are also key concerns. Collaborating closely with software architects and data scientists is common, as solutions frequently require balancing algorithmic accuracy with hardware limitations. Addressing these challenges requires continuous profiling, testing, and iterative optimization.

What is a CUDA Kernel Engineer?

Cuda Kernel Engineers are specialized software developers who design, implement, and optimize parallel computing algorithms using NVIDIA's CUDA platform. They write 'kernels,' which are functions that run on Graphics Processing Units (GPUs) to accelerate computational tasks in areas such as machine learning, scientific simulations, and graphics rendering. These engineers need strong skills in C/C++ programming, GPU architecture, and performance optimization techniques. Their work is crucial for applications that require high-speed data processing and efficient resource utilization.

What skills and qualifications are needed to be a CUDA Kernel Engineer?

To thrive as a CUDA Kernel Engineer, you need strong proficiency in C/C++ programming, parallel computing concepts, and a solid foundation in GPU architectures, typically supported by a degree in computer science or a related field. Expertise in NVIDIA CUDA toolkits, GPU profiling tools like Nsight, and familiarity with version control systems are essential. Analytical thinking, problem-solving abilities, and effective collaboration skills help engineers optimize code and work well within development teams. These skills and qualities are crucial for delivering high-performance, scalable GPU solutions in computationally intensive applications.
What job categories do people searching Cuda Kernel Engineer jobs in West Roxbury, MA look for? The top searched job categories for Cuda Kernel Engineer jobs in West Roxbury, MA are:
What cities near West Roxbury, MA are hiring for Cuda Kernel Engineer jobs? Cities near West Roxbury, MA with the most Cuda Kernel Engineer job openings:

Senior Robotics Middleware Engineer

Integrated Computer Solutions, Inc.

Wilmington, MA • On-site

Contractor

Posted 15 days ago


Job description

Location/Eligibility: This is a contract position based in Massachusetts.  You must be able to commute to the Wilmington area onsite at least 3 days a week.  For this contract, there is no visa sponsorship available.  

We are seeking a high-caliber Robotics Middleware Engineer to evaluate, prototype, and build our next-generation communication infrastructure. In this role, you will move beneath the standard application layer to optimize high-throughput, shared-memory messaging on NVIDIA Jetson-class hardware. You will be the architect responsible for eliminating communication bottlenecks, ensuring that heavy ML inference and sensor streams flow through the robot's "nervous system" with zero latency.

Core Responsibilities
  • Next-Gen Middleware Evaluation: Lead a comprehensive build-vs-adopt evaluation of Zenoh as a robotics transport layer, analyzing its performance both standalone and as an RMW layer under ROS 2. Deliver architectural recommendations backed by rigorous benchmarks.

  • Data Path Optimization: Optimize zero-copy and shared-memory messaging paths on ARM/Jetson architectures, focusing heavily on memory layout optimization, serialization efficiency, and transport tuning.

  • Performance Benchmarking: Design and execute benchmarking suites to measure latency, throughput, and CPU/GPU overhead, systematically comparing new architectures against our current data path.

  • ML & Sensor Integration: Support the full-stack team by ensuring low-latency integration for heavy workloads, including vision sensors (Stereo/RGBD), IMUs, and ML model inference pipelines (PyTorch, TensorRT, JIT).

Technical Requirements (Must-Have)
  • ROS 2 & DDS Internals: Deep, hands-on experience with ROS 2 middleware (RMW) layers and the inner workings of DDS implementations (e.g., Cyclone DDS, Fast DDS).

  • Advanced Transports: Proven experience with Zenoh or comparable modern pub/sub architectures designed for high-performance edge computing (e.g., eCAL, iceoryx).

  • Systems Programming: Exceptional proficiency in Modern C++ and deep familiarity with shared-memory Inter-Process Communication (IPC).

  • Embedded Profiling: Hands-on experience with performance profiling and bottleneck identification on embedded Linux and ARM architectures.

  • Core Languages: Strong proficiency in Python is required alongside your primary C++ systems skill set.

Preferred Qualifications (Nice-to-Have)
  • NVIDIA Jetson Ecosystem: Direct experience optimizing software for Jetson-class hardware.

  • Hardware-Aware Memory: Understanding of GPU/CUDA memory awareness, unified memory architectures, and avoiding host-to-device copy overhead.

  • Real-Time Systems: Practical knowledge of real-time Linux patches (PREEMPT_RT) and deterministic software design.

  • Experiencing tuning or customizing the ROS 2 Nav2 stack or working with vSLAM algorithms.

Are you the right fit?

You thrive in high-velocity environments where engineering decisions directly impact hardware performance. You enjoy digging into network packets, memory dumps, and kernel-level timing to squeeze microsecond efficiencies out of resource-constrained systems. You value data over assumptions and back up your architectural designs with hard benchmarks.