Understanding of GPU/CUDA memory awareness, unified memory architectures, and avoiding host-to ... You enjoy digging into network packets, memory dumps, and kernel-level timing to squeeze ...
Quick apply
Understanding of GPU/CUDA memory awareness, unified memory architectures, and avoiding host-to ... You enjoy digging into network packets, memory dumps, and kernel-level timing to squeeze ...
Quick apply
Understanding of GPU/CUDA memory awareness, unified memory architectures, and avoiding host-to ... You enjoy digging into network packets, memory dumps, and kernel-level timing to squeeze ...
Training, Distributed inference, MoE, Reinforcement Learning, kernel authoring (on CUDA, Triton, cuTe, etc). Experience with programming for compute & communication overlap in distributed runtimes
Training, Distributed inference, MoE, Reinforcement Learning, kernel authoring (on CUDA, Triton, cuTe, etc). Experience with programming for compute & communication overlap in distributed runtimes
Boston, MA · On-site +1
$144K - $192K/yr
Kernel Development : Design and maintain high-performance GPU kernels in Triton or CUDA for state-of-the-art ML workloads. * Data Pipeline Engineering : Optimize robust data loading pipelines that ...
Quick apply
Boston, MA · On-site +1
$144K - $192K/yr
Kernel Development : Design and maintain high-performance GPU kernels in Triton or CUDA for state-of-the-art ML workloads. * Data Pipeline Engineering : Optimize robust data loading pipelines that ...
Boston, MA · On-site +1
$144K - $192K/yr
Kernel Development : Design and maintain high-performance GPU kernels in Triton or CUDA for state-of-the-art ML workloads. * Data Pipeline Engineering : Optimize robust data loading pipelines that ...
Boston, MA · On-site +1
$144K - $192K/yr
Kernel Development : Design and maintain high-performance GPU kernels in Triton or CUDA for state-of-the-art ML workloads. * Data Pipeline Engineering : Optimize robust data loading pipelines that ...
Responsibilities include profiling, performance optimization, custom kernel development ... C++/CUDA a plus * Experience with distributed ML training frameworks (Megatron-LM, TorchTitan ...
Responsibilities include profiling, performance optimization, custom kernel development ... C++/CUDA a plus * Experience with distributed ML training frameworks (Megatron-LM, TorchTitan ...
$60K - $85K/yr
Experience with CUDA programming; Experience programming distributed systems; Experience with ... Advanced knowledge of Linux kernel internals and systems programming methods; Advanced knowledge of ...
$60K - $85K/yr
Experience with CUDA programming; Experience programming distributed systems; Experience with ... Advanced knowledge of Linux kernel internals and systems programming methods; Advanced knowledge of ...
Boston, MA · On-site
$60K - $85K/yr
Experience with CUDA programming; Experience programming distributed systems; Experience with ... Advanced knowledge of Linux kernel internals and systems programming methods; Advanced knowledge of ...
Boston, MA · On-site
$60K - $85K/yr
Experience with CUDA programming; Experience programming distributed systems; Experience with ... Advanced knowledge of Linux kernel internals and systems programming methods; Advanced knowledge of ...
We need an engineer to develop and build an automated framework. This framework will ingest ... Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory ...
We need an engineer to develop and build an automated framework. This framework will ingest ... Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory ...
$127K - $167K/yr
Hands-on experience with GPU kernel development or optimization (CUDA/C++, Triton, or equivalent ... Container engineering expertise: multi-architecture Docker or OCI builds, layer optimization ...
$127K - $167K/yr
Hands-on experience with GPU kernel development or optimization (CUDA/C++, Triton, or equivalent ... Container engineering expertise: multi-architecture Docker or OCI builds, layer optimization ...
Cambridge, MA · On-site
$189K - $289K/yr
Responsibilities include profiling, performance optimization, custom kernel development ... C++/CUDA a plus * Experience with distributed ML training frameworks (Megatron-LM, TorchTitan ...
Cambridge, MA · On-site
$189K - $289K/yr
Responsibilities include profiling, performance optimization, custom kernel development ... C++/CUDA a plus * Experience with distributed ML training frameworks (Megatron-LM, TorchTitan ...
Boston, MA · On-site
$120K - $170K/yr
... kernel schedulers. * 3+ years of startup experience - you know the trade-offs between shipping fast ... CUDA and/or Jetson platform * Experience with classical computer vision techniques and machine ...
Boston, MA · On-site
$120K - $170K/yr
... kernel schedulers. * 3+ years of startup experience - you know the trade-offs between shipping fast ... CUDA and/or Jetson platform * Experience with classical computer vision techniques and machine ...
Boston, MA · On-site
$120K - $170K/yr
... kernel schedulers. * 3+ years of startup experience - you know the trade-offs between shipping fast ... CUDA and/or Jetson platform * Experience with classical computer vision techniques and machine ...
Boston, MA · On-site
$120K - $170K/yr
... kernel schedulers. * 3+ years of startup experience - you know the trade-offs between shipping fast ... CUDA and/or Jetson platform * Experience with classical computer vision techniques and machine ...
$120K - $170K/yr
... kernel schedulers. * 3+ years of startup experience - you know the trade-offs between shipping fast ... CUDA and/or Jetson platform * Experience with classical computer vision techniques and machine ...
Quick apply
$120K - $170K/yr
... kernel schedulers. * 3+ years of startup experience - you know the trade-offs between shipping fast ... CUDA and/or Jetson platform * Experience with classical computer vision techniques and machine ...
A grasp of the CUDA programming model and experience employing GPU profiling tools like NVIDIA Nsight Systems/Compute to address PCIe bottlenecks and kernel stalls. * Extensive knowledge of profiling ...
A grasp of the CUDA programming model and experience employing GPU profiling tools like NVIDIA Nsight Systems/Compute to address PCIe bottlenecks and kernel stalls. * Extensive knowledge of profiling ...
Collaborate with ML and software engineering colleagues to deploy and operationalize models ... Training and inference optimization (e.g., mixed precision, kernel optimization, quantization ...
Collaborate with ML and software engineering colleagues to deploy and operationalize models ... Training and inference optimization (e.g., mixed precision, kernel optimization, quantization ...
Contractor
Posted 15 days ago
We are seeking a high-caliber Robotics Middleware Engineer to evaluate, prototype, and build our next-generation communication infrastructure. In this role, you will move beneath the standard application layer to optimize high-throughput, shared-memory messaging on NVIDIA Jetson-class hardware. You will be the architect responsible for eliminating communication bottlenecks, ensuring that heavy ML inference and sensor streams flow through the robot's "nervous system" with zero latency.
Core ResponsibilitiesNext-Gen Middleware Evaluation: Lead a comprehensive build-vs-adopt evaluation of Zenoh as a robotics transport layer, analyzing its performance both standalone and as an RMW layer under ROS 2. Deliver architectural recommendations backed by rigorous benchmarks.
Data Path Optimization: Optimize zero-copy and shared-memory messaging paths on ARM/Jetson architectures, focusing heavily on memory layout optimization, serialization efficiency, and transport tuning.
Performance Benchmarking: Design and execute benchmarking suites to measure latency, throughput, and CPU/GPU overhead, systematically comparing new architectures against our current data path.
ML & Sensor Integration: Support the full-stack team by ensuring low-latency integration for heavy workloads, including vision sensors (Stereo/RGBD), IMUs, and ML model inference pipelines (PyTorch, TensorRT, JIT).
ROS 2 & DDS Internals: Deep, hands-on experience with ROS 2 middleware (RMW) layers and the inner workings of DDS implementations (e.g., Cyclone DDS, Fast DDS).
Advanced Transports: Proven experience with Zenoh or comparable modern pub/sub architectures designed for high-performance edge computing (e.g., eCAL, iceoryx).
Systems Programming: Exceptional proficiency in Modern C++ and deep familiarity with shared-memory Inter-Process Communication (IPC).
Embedded Profiling: Hands-on experience with performance profiling and bottleneck identification on embedded Linux and ARM architectures.
Core Languages: Strong proficiency in Python is required alongside your primary C++ systems skill set.
NVIDIA Jetson Ecosystem: Direct experience optimizing software for Jetson-class hardware.
Hardware-Aware Memory: Understanding of GPU/CUDA memory awareness, unified memory architectures, and avoiding host-to-device copy overhead.
Real-Time Systems: Practical knowledge of real-time Linux patches (PREEMPT_RT) and deterministic software design.
Experiencing tuning or customizing the ROS 2 Nav2 stack or working with vSLAM algorithms.
You thrive in high-velocity environments where engineering decisions directly impact hardware performance. You enjoy digging into network packets, memory dumps, and kernel-level timing to squeeze microsecond efficiencies out of resource-constrained systems. You value data over assumptions and back up your architectural designs with hard benchmarks.
Sourced by ZipRecruiter
Software development
51 - 200 Employees
Waltham, MA, US
1987