1

Customer Reliability Engineer Jobs in Micronesia

Collaborate with architecture, hardware, firmware, software, and Customer engagement teams to improve system quality and reliability. * Build and maintain SRE-style validation infrastructure ...

... level of asset reliability possible. Our solutions for pre-commissioning, commissioning, and ... Travel to customer sites to provide supporting service. * Completing timely reports that identify ...

... level of asset reliability possible. Our solutions for pre-commissioning, commissioning, and ... Travel to customer sites to provide supporting service. * Completing timely reports that identify ...

$108K - $150K/yr

... and customer expectations. * Establish and promote engineering best practices to improve product performance, reliability, and time-to-market (TTM). * Develop, review, and maintain technical ...

$94K - $130K/yr

... customer requirements. * Design and develop 3U and 6U VPX SBC, FPGA, switch, and I/O cards ... Apply engineering best practices to improve product performance, reliability, and time-to-market ...

$94K - $130K/yr

... customer requirements. * Design and develop 3U and 6U VPX SBC, FPGA, switch, and I/O cards ... Apply engineering best practices to improve product performance, reliability, and time-to-market ...

... customers, and our culture. A Little About You You bring a unique blend of personality and ... Ensure application performance, scalability, security, and reliability through best engineering ...

$139K - $170K/yr

Some travel to integration sites, customer facilities, and subcontractor facilities may be required ... Review of electronic component and assembly reliability and construction data, and preparation of ...

$144K - $190K/yr

Support business development, program managers, and customer teams in translating requirements into ... performance, reliability, and time-to-market (TTM). * Direct and conduct tests to establish ...

$144K - $190K/yr

Support business development, program managers, and customer teams in translating requirements into ... performance, reliability, and time-to-market (TTM). * Direct and conduct tests to establish ...

$122K - $161K/yr

Work with internal customers to develop application requirements. * Understand and follow existing ... Create monitoring and alerting solutions for system health and reliability. Cloud & DevOps * Deploy ...

$104K - $143K/yr

Critical Design Review(CDR), Manufacturing Readiness Assessment (MRA), Reliability Technical ... Facilitate the transfer of information, lessons learned and best practices across all customers and ...

$104K - $143K/yr

... customers including data analysis, modification programs, repair requirements, reliability ... Provide engineering review of programs throughout the life-cycle * Review and ensure all ...

Partner directly with hyperscalers, CSPs, enterprise customers, internal component leads, vendors ... Establish reliability, security, validation, and left-shift strategies that reduce risk before ...

$253K/yr

... customer-centric solutions * Drive adoption of modern engineering practices and technologies to improve speed, quality, and reliability * Ensure strong alignment across platform and service teams to ...

$253K/yr

... customer-centric solutions * Drive adoption of modern engineering practices and technologies to improve speed, quality, and reliability * Ensure strong alignment across platform and service teams to ...

$49 - $63/hr

... and reliability. As a COBOL Mainframe Application Developer, you will work directly with the ... You prioritize customer success and are dedicated to delivering exceptional results. *Candidates ...

... customer priorities Test, refine and automate processes to improve reliability, efficiency and ... Engineer roles Kyndryl currently does not require employees to be fully vaccinated against COVID-19 ...

Summary: The Manager, Facilities Operations & Engineering is a front-line operations leader ... Responds in a timely manner to all customer complaints, requests for service or general inquiries.

Customer Reliability Engineer information

What is a customer reliability engineer?

A Customer Reliability Engineer (CRE) is a technical professional who works closely with customers to ensure the reliability, performance, and uptime of software products and services. CREs act as a bridge between customers and engineering teams, helping to identify, troubleshoot, and resolve reliability issues. They often collaborate with multiple departments to implement best practices, monitor systems, and proactively address potential problems, ultimately aiming to improve the overall customer experience.

What skills and qualifications are needed to thrive as a customer reliability engineer?

To thrive as a Customer Reliability Engineer, you need a solid background in systems engineering, incident management, and troubleshooting, often supported by a degree in computer science or related field. Familiarity with cloud platforms (such as AWS or GCP), monitoring tools (like Datadog or Prometheus), and automation scripts is typically required. Exceptional communication, problem-solving abilities, and a customer-centric mindset are vital soft skills for this role. These skills ensure efficient incident resolution, strong client relationships, and reliable system performance under pressure.

How does a customer reliability engineer typically interact with clients and internal engineering teams?

Customer Reliability Engineers serve as a vital bridge between clients and internal technical teams. They regularly communicate with customers to understand their needs, troubleshoot issues, and provide technical guidance. Internally, they collaborate closely with product, support, and development teams to relay customer feedback, help prioritize reliability improvements, and ensure seamless incident resolution. This cross-functional role requires strong communication skills and the ability to translate technical information for different audiences, making every day varied and impactful.

What is the difference between Customer Reliability Engineer vs Site Reliability Engineer?

AspectCustomer Reliability EngineerSite Reliability Engineer
CredentialsTypically requires engineering degrees, certifications in cloud platforms (AWS, Azure), and knowledge of customer supportRequires engineering degrees, certifications in cloud and systems management, with a focus on infrastructure
Work EnvironmentCustomer-facing, involves direct interaction with clients to resolve issues and improve reliabilityPrimarily internal, focused on maintaining and improving system reliability and scalability
Employer & Industry UsageUsed by cloud service providers and tech companies with a customer support componentCommon in large tech companies managing large-scale infrastructure and services

The main difference is that Customer Reliability Engineers focus on ensuring customer satisfaction and resolving client-specific issues, while Site Reliability Engineers concentrate on internal system stability and scalability. Both roles require technical expertise and cloud knowledge but serve different operational needs.

How much do customer reliability engineers get paid?

Customer Reliability Engineers typically earn a median annual salary between $80,000 and $130,000, depending on experience, location, and industry. They often require skills in cloud platforms, troubleshooting, and automation tools, which can influence compensation levels.

What does a customer reliability engineer do?

A customer reliability engineer (CRE) is responsible for ensuring the reliability, performance, and availability of a company's products or services for customers. They analyze system issues, collaborate with development and support teams, and implement solutions to improve customer experience, often using monitoring tools and technical troubleshooting skills.
Infographic showing various Customer Reliability Engineer job openings in Micronesia as of August 2026, with employment types broken down into 79% Full Time, 19% Part Time, and 2% Contract. Highlights an 88% Physical, 1% Hybrid, and 11% Remote job distribution.

Senior Software Engineer - NVLink Rack Scale Stability and Reliability

Nvidia

On-site, Remote

Full-time

Re-posted 11 days ago


Nvidia rating

9.6

Company rating: 9.6 out of 10

Based on 17 frontline employees who took The Breakroom Quiz

7th of 245 rated software companies


Job description

NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is looking for phenomenal people like you to help us accelerate the next wave of artificial intelligence.

We are looking for highly motivated Senior Software Engineers to join our Fabric Networking team with a targeted focus on NVLink Rack-Scale Systems Stability & Reliability. In this role, you will partner closely with architects and developers building our next-generation NVLink and NVSwitch systems, helping transform first-of-their-kind platforms into stable, reliable, and volume production-ready systems. You will work on complex system-level challenges spanning resiliency, diagnostics, recovery, and large-scale AI infrastructure, contributing directly to the software foundation powering next-generation datacenter deployments.

What you will be doing:

  • Drive platform bringup, feature enablement, end-to-end software validation, and debug for next-generation NVLink-based GPU and rack-scale systems.

  • Develop tools, diagnostics, automation, and infrastructure for system validation, regression testing, and fleet support.

  • Lead reliability and MTBI validation through stress testing, telemetry analysis, failure injection, and issue resolution.

  • Triage complex software, firmware, networking, and platform issues across validation, deployment, and production environments.

  • Collaborate with architecture, hardware, firmware, software, and Customer engagement teams to improve system quality and reliability.

  • Build and maintain SRE-style validation infrastructure, including provisioning, monitoring, and operational readiness.

  • Create automation, dashboards, runbooks, and debug workflows that improve root-cause analysis and operational efficiency.

What we need to see:

  • BS or MS in Computer Science, Computer Engineering, Electrical Engineering, or related field, or equivalent experience.

  • 5+ years of experience in system software, firmware, networking, platform enablement, data center infrastructure, or distributed systems.

  • Strong programming skills in C/C++ and Python; Bash/Shell scripting experience is a plus.

  • Strong system-level debugging across software, firmware, hardware, and networking layers.

  • Solid networking fundamentals, including TCP/IP, Ethernet and/or InfiniBand, RDMA/RoCE, routing, switching, and fabric performance analysis.

  • Experience with large-scale AI systems, including platform bringup, validation, reliability engineering, stress testing, telemetry analysis, and root-cause debugging.

  • Ability to triage complex multi-domain issues using logs, telemetry, experiments, and structured debugging methods.

  • Strong communication and collaboration skills across engineering, customer, and operations teams.Passion for building reliable next-generation AI infrastructure and solving complex system-level challenges at scale.

Ways to stand out from the crowd:

  • Experience with NVIDIA GPU systems, NVLink, NVSwitch, CUDA, and large-scale AI/HPC clusters such as NVIDIA GB200 NVL72.

  • Strong understanding of large-scale AI system architecture, including PCIe, memory hierarchy, DMA, high-speed interconnects, and distributed training/inference systems.

  • Experience with server management technologies, data center operations, cluster provisioning, scaling, and fleet monitoring.

  • Proven experience building diagnostics, automation, CI/CD pipelines, dashboards, and reliability tooling.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 21, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

What Nvidia employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom


Nvidia logo

About Nvidia

Sourced by ZipRecruiter

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology--and amazing people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what's never been done before takes vision, innovation, and the world's best talent.

Industry

Computer and electronic product manufacturing

Company size

10,000+ Employees

Headquarters location

Santa Clara, CA, US

Year founded

1993