2

Remote Gpu Engineer Jobs in Missouri (NOW HIRING)

Build comprehensive observability across compute, storage, networking, GPU clusters, inference ... Permanent, full-time position with an EMEA-based remote work model. * Flexible and hybrid-friendly ...

$94K - $124K/yr

You will have significant autonomy to shape APIs, streaming protocols, GPU serving, observability ... Flexible working hours and remote work options . * Health, dental, and vision benefits for ...

New

$100K - $180K/yr

The environment is remote-first, highly collaborative, and designed for people who take ownership ... GPU server and related infrastructure issues. * Capture and communicate customer feedback and ...

Optimize CPU, GPU, memory, battery usage, and overall performance across a broad range of mobile ... Remote-first and flexible working environment. * Digital nomad-friendly culture. * Generous paid ...

Experience operating multi-tenant AI platforms or managing GPU and AI resources at scale is ... Fully remote working environment with the flexibility to work from anywhere. * One-time home office ...

$88K - $121K/yr

This is a senior security engineering opportunity focused on protecting a high-performance GPU ... Remote work opportunity across Europe with a flexible, hybrid-friendly working approach. * Friendly ...

You will work at the intersection of AI research, model architecture, systems engineering, and ... Implement systems-level optimizations including dynamic batching, kernel fusion, multi-GPU ...

... engineering teams. Benefits: * Full-time opportunity within a research-focused AI environment. * Remote working arrangement with the flexibility to contribute from locations worldwide. * Direct ...

New

Remote Gpu Engineer information

What is a remote GPU engineer?

Remote GPU Engineers are specialized software or hardware engineers who work primarily with Graphics Processing Units (GPUs) from a remote location. They focus on designing, optimizing, and maintaining GPU-based systems for applications such as machine learning, high-performance computing, and graphics rendering. These professionals often collaborate with teams virtually, leveraging cloud-based GPU resources and remote access tools. Their work enables companies to efficiently utilize GPU technology without requiring engineers to be on-site.

What are the key skills and qualifications needed to thrive as a remote GPU engineer?

To thrive as a Remote GPU Engineer, you need strong expertise in GPU architectures, parallel programming (CUDA/OpenCL), and a solid background in computer science or engineering. Familiarity with tools like CUDA Toolkit, performance profilers, and version control systems, as well as experience with relevant certifications, is typically required. Excellent problem-solving abilities, communication skills, and the capacity to collaborate effectively in remote, distributed teams are standout soft skills. These competencies ensure efficient GPU solution development, effective troubleshooting, and seamless teamwork in a remote engineering environment.

What are some common challenges faced by remote GPU engineers when collaborating with distributed teams?

Remote GPU Engineers often work with global teams, which can present challenges such as coordinating across different time zones, ensuring consistent communication, and managing access to high-performance hardware remotely. To overcome these hurdles, it's important to leverage collaboration tools, maintain clear documentation, and establish regular check-ins. Additionally, using remote desktop solutions and cloud-based GPU environments can help facilitate smoother development and debugging processes.

What are the most commonly searched types of Gpu Engineer jobs in Missouri?

The most popular types of Gpu Engineer jobs in Missouri are:

What are popular job titles related to Remote Gpu Engineer jobs in Missouri?

For Remote Gpu Engineer jobs in Missouri, the most frequently searched job titles are:

What cities in Missouri are hiring for Remote Gpu Engineer jobs?

Cities in Missouri with the most Remote Gpu Engineer job openings:

Senior Observability & Telemetry Engineer - Radian Arc

Jobgether

Remote

Full-time

Posted 10 days ago


Key responsibilities

  • Design, implement, and operate scalable telemetry pipelines for metrics, logs, and traces across distributed GPU and edge infrastructure.

  • Architect and maintain telemetry storage systems optimized for large-scale time-series and event data, and contribute to standards for instrumentation, logging, tracing, and SLOs.

  • Develop dashboards, monitoring tools, and performance-analysis capabilities to provide insights into workload health, GPU utilization, storage throughput, network latency, and inference performance.


Job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Observability & Telemetry Engineer - Radian Arc based in Netherlands.

This is a high-impact engineering role focused on building the observability foundation for large-scale GPU cloud and edge infrastructure.
You will design and operate telemetry platforms that provide real-time visibility across distributed AI workloads, compute, storage, networking, and inference environments.
The role combines observability architecture, infrastructure telemetry, reliability engineering, and customer-facing performance insights.
You will work with high-cardinality metrics, logs, traces, and event data across both hyperscale environments and smaller edge deployments.
Your work will directly improve platform reliability, operational efficiency, performance, and incident response.
As a senior engineer, you will lead major initiatives, influence observability standards, and mentor engineers across multiple technical teams.
The environment is international, technically ambitious, and well suited to someone who enjoys solving complex distributed-systems challenges with significant autonomy.

Accountabilities
  • Design, implement, and operate scalable telemetry pipelines for metrics, logs, and traces across distributed GPU and edge infrastructure, supporting high-cardinality data from thousands of nodes and services.
  • Architect and maintain telemetry storage systems optimized for large-scale time-series and event data, while contributing to standards for instrumentation, logging, tracing, and SLO implementation.
  • Build comprehensive observability across compute, storage, networking, GPU clusters, inference workloads, and distributed training environments, identifying issues such as GPU throttling, network congestion, storage latency, and hardware degradation.
  • Develop dashboards, monitoring tools, and performance-analysis capabilities that provide internal teams and customers with actionable insights into workload health, GPU utilization, storage throughput, network latency, and inference performance.
  • Build and maintain network and infrastructure telemetry solutions using Python or Go, integrating data from technologies such as NVIDIA Cumulus Linux, VyOS, Citrix NetScaler/WAF, gNMI, SNMP, and streaming telemetry.
  • Develop advanced alerting, anomaly detection, reliability metrics, SLIs, and SLOs, integrating observability signals into operational workflows and incident management processes.
  • Collaborate with platform, networking, storage, compute, and operations teams to improve instrumentation, monitoring, incident response, and platform reliability.
  • Provide technical guidance and mentorship to engineers, promoting effective observability practices and consistent monitoring patterns across the organization.
  • Participate in on-call rotations supporting production observability and telemetry infrastructure.
Requirements
  • Proven experience operating observability systems and distributed infrastructure platforms at production scale, with strong expertise across metrics, logging, tracing, alerting, dashboards, and telemetry pipelines.
  • Strong programming skills in Go, Python, or Rust, with experience developing telemetry collectors, exporters, automation, or infrastructure tooling.
  • Hands-on experience with observability technologies such as Prometheus, OpenTelemetry, Grafana, distributed logging platforms, and large-scale telemetry databases such as ClickHouse or equivalent.
  • Experience working with large-scale GPU cloud, HPC, or AI infrastructure and monitoring distributed training or inference workloads.
  • Knowledge of GPU telemetry technologies such as NVIDIA DCGM, DCGM Exporter, NVML, GPU Operator telemetry, NVLink, and NVSwitch, as well as AI workload metrics including inference latency, throughput, NCCL health, synchronization latency, and storage I/O.
  • Strong understanding of networking and infrastructure telemetry, including experience with gNMI, SNMP, streaming telemetry, network flow telemetry, RDMA/RoCE, or comparable technologies.
  • Familiarity with cloud-native infrastructure, including Kubernetes, automation, CI/CD, and distributed systems.
  • Strong analytical and troubleshooting capabilities, with the ability to interpret complex telemetry signals, diagnose performance problems, identify systemic issues, and translate findings into actionable improvements.
  • Excellent collaboration and communication skills, with the ability to work effectively across infrastructure, networking, storage, compute, and operations teams.
  • A proactive, ownership-oriented mindset and the ability to lead complex observability initiatives in a fast-moving, technically sophisticated environment.
Benefits
  • Attractive compensation package reflecting your expertise, experience, transferable skills, and market conditions.
  • Permanent, full-time position with an EMEA-based remote work model.
  • Flexible and hybrid-friendly working environment designed to support international collaboration.
  • Opportunity to work on large-scale GPU, AI, cloud, networking, and edge infrastructure challenges.
  • Exposure to advanced observability, telemetry, reliability, and distributed-systems technologies.
  • Opportunity to contribute to major platform initiatives and influence observability standards across multiple engineering teams.
  • Mentorship and leadership opportunities, including the ability to guide engineers and promote best practices.
  • Career growth within a fast-growing international scale-up focused on innovative infrastructure solutions.
  • Inclusive and diverse working environment where qualified candidates are considered fairly and supported in their development.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
 Why Apply Through Jobgether? 
 
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
 
 
#LI-CL1
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
apply for this job