2

Remote Observability Engineer Jobs in Renton, WA

Senior Software Engineer

Seattle, WA ยท On-site +1

$164K - $266K/yr

What you'll do As a Senior Software Engineer on the Observability team, you will design, build, and ... Employee divides their time between in-office and remote work. Access to an office location is ...

... observability best practices, product usability, and SRE standards. Your Impact * Experienced operational leader who understands incident response, customer critical issue dynamics, engineering ...

... observability, and security. You collaborate well in cross-functional settings and help raise the ... Employee divides their time between in-office and remote work. Access to an office location is ...

Lead DevOps Engineer

Seattle, WA ยท Remote

$54 - $74/hr

Drive adoption of observability tools like DataDog and establish logging standards * Coordinate ... Remote * Contract or B2B arrangement Our values We are a company that seeks the best for both our ...

Senior Software Engineer - AI Platform

Seattle, WA ยท Remote

$139K - $183K/yr

This is a remote position; however, the candidate must reside within 30 miles of one of the ... Contribute to and refine platform standards for security, compliance, observability, and resiliency ...

Senior Software Engineer - AI Platform

Seattle, WA ยท Remote

$139K - $183K/yr

This is a remote position; however, the candidate must reside within 30 miles of one of the ... Contribute to and refine platform standards for security, compliance, observability, and resiliency ...

Senior Software Engineer - AI Platform

Seattle, WA ยท Remote

$139K - $183K/yr

This is a remote position; however, the candidate must reside within 30 miles of one of the ... Contribute to and refine platform standards for security, compliance, observability, and resiliency ...

Senior Software Engineer

Seattle, WA ยท On-site +1

$139K - $183K/yr

Implement high-quality, testable, and production-ready code with strong observability, including ... Employee divides their time between in-office and remote work. Access to an office location is ...

next page

Showing results 1-20

Remote Observability Engineer information

See Renton, WA salary details

$42.7K

$130.3K

$215.4K

How much do remote observability engineer jobs pay per year?

As of Jun 19, 2026, the average yearly pay for remote observability engineer in Renton, WA is $130,327.00, according to ZipRecruiter salary data. Most workers in this role earn between $93,400.00 and $170,400.00 per year, depending on experience, location, and employer.

What are the typical collaboration patterns for a Remote Observability Engineer working with distributed teams?

Remote Observability Engineers frequently collaborate with software developers, DevOps teams, and IT operations to ensure systems are monitored effectively and issues are detected early. Working remotely, you'll often use communication tools like Slack, Jira, and video conferencing to coordinate incident response, discuss monitoring strategies, and review system health dashboards. Regular sync meetings and asynchronous updates are common, and you'll likely contribute to documentation and knowledge sharing to keep all stakeholders informed. Building strong communication habits is important, as much of the troubleshooting and improvement work hinges on clear coordination with multiple teams.

What are the key skills and qualifications needed to thrive as a Remote Observability Engineer, and why are they important?

To thrive as a Remote Observability Engineer, you need strong expertise in monitoring, logging, and tracing systems, along with a background in computer science or related technical fields. Familiarity with tools like Prometheus, Grafana, ELK Stack, Datadog, and cloud platforms is typically required, as well as relevant certifications such as AWS Certified Cloud Practitioner or Google Cloud Professional DevOps Engineer. Excellent problem-solving abilities, communication skills, and a proactive mindset help you detect and resolve issues before they impact users. These competencies ensure system reliability, enable rapid incident response, and support seamless collaboration in distributed environments.

What is the difference between Remote Observability Engineer vs Site Reliability Engineer?

AspectRemote Observability EngineerSite Reliability Engineer
CredentialsKnowledge of monitoring tools, scripting, cloud platformsSame as Observability Engineer, plus SRE certifications often preferred
Work EnvironmentFocus on monitoring, logging, and tracing systems remotelyBroader scope including system reliability, incident response, and automation
Industry UsagePrimarily in tech, SaaS, cloud servicesWidely in tech, finance, and large-scale online services

The Remote Observability Engineer specializes in monitoring and analyzing system performance remotely, focusing on tools like logs and metrics. In contrast, the Site Reliability Engineer has a broader role, ensuring overall system reliability, automation, and incident management. While both roles require similar technical skills, SREs often have additional responsibilities related to system resilience and scalability.

What is a Remote Observability Engineer?

A Remote Observability Engineer is a professional responsible for designing, implementing, and maintaining systems that monitor the health, performance, and reliability of software applications and infrastructure from a remote location. They use observability tools to collect and analyze logs, metrics, and traces, helping organizations quickly detect and resolve issues. Their work ensures that distributed systems are transparent, reliable, and efficient, often collaborating with development, operations, and security teams. Remote Observability Engineers often work from anywhere, leveraging cloud-based tools and platforms to manage complex IT environments.
What are popular job titles related to Remote Observability Engineer jobs in Renton, WA? For Remote Observability Engineer jobs in Renton, WA, the most frequently searched job titles are:
What job categories do people searching Remote Observability Engineer jobs in Renton, WA look for? The top searched job categories for Remote Observability Engineer jobs in Renton, WA are:
What cities near Renton, WA are hiring for Remote Observability Engineer jobs? Cities near Renton, WA with the most Remote Observability Engineer job openings:
Infrastructure Engineer (Observability)

Infrastructure Engineer (Observability)

Lightning AI

Seattle, WA โ€ข On-site, Remote

$122K - $160K/yr

Full-time

Medical, Dental, Vision, Retirement, PTO

Posted 8 days ago


Job description

Who We Are

Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systemsโ€”designed to take ideas from research to production with less friction.

Through our merger with Voltage Park, a neocloud and AI Factory, Lightning AI combines developer-first software with cost-efficient, large-scale compute. Teams get the tools they need for experimentation, training, and production inference, with security, observability, and control built in.

We serve solo researchers, startups, and large enterprises. Lightning AI operates globally with offices in New York City, San Francisco, Seattle, and London, and is backed by Coatue, Index Ventures, Bain Capital Ventures, and Firstminute.

ย What Weโ€™re Looking For

Lightning AI is seeking an Observability Infrastructure Engineer to join our Infrastructure Engineering team.

In this role, you will own and evolve observability systems across large-scale, GPU-enabled bare-metal infrastructure. Youโ€™ll operate at the intersection of infrastructure, data, and product, building platforms for metrics, logs, traces, and alerting that power both internal operations and customer-facing visibility.

You will play a key role in productizing observability, enabling scalable, multi-tenant monitoring experiences while keeping pace with rapid infrastructure buildouts. This includes designing telemetry pipelines, improving signal quality, and delivering actionable insights that ensure reliability and transparency across our platform.

Weโ€™re flexible on location for this team. This role can work hybrid out of one of our US-based hubs (Seattle, NYC, or SF) or fully remote within the U.S., with occasional company and team offsites. We are not able to provide visa sponsorship for this position at this time.

What Youโ€™ll Do Observability Platform & Productization
  • Own and evolve a scalable observability platform spanning metrics, logs, traces, and events
  • Drive the productization of observability capabilities for both internal teams and external customers
  • Design multi-tenant observability systems with scoped access, RBAC, and customer-facing visibility
  • Continuously improve observability systems to keep pace with rapid infrastructure buildouts
Telemetry & Data Pipelines
  • Design and operate telemetry pipelines ingesting data from GPUs, CPUs, networking (Ethernet & InfiniBand), containers, APIs, and BMC/Redfish
  • Build systems to correlate signals across infrastructure layers to enable faster debugging and root cause analysis
  • Implement streaming and real-time data pipelines using tools such as Kafka, OTEL, Promtail, or similar
Alerting, Reliability & Insights
  • Design and implement noise-resistant alerting systems to improve signal quality and reduce operational load
  • Create dashboards and alerting for InfraOps, Engineering, and Customer Success teams
  • Build automated insights and enable proactive detection, forecasting, and system health visibility at scale
Systems & Infrastructure Engineering
  • Contribute to broader infrastructure engineering projects beyond observability
  • Partner with infrastructure and platform teams to embed observability into core systems and workflows
  • Support large-scale, distributed systems across compute, networking, and storage environments
Cross-Functional Collaboration
  • Work closely with customer-facing teams to deliver external observability experiences
  • Collaborate with engineering, operations, and support teams to improve system transparency and reliability
  • Help define best practices for observability across the organization
What Youโ€™ll Need Required Qualifications
  • 5+ years of experience in infrastructure engineering, SRE, or observability-focused roles
  • Strong experience with monitoring systems such as Prometheus, Grafana, ELK, or VictoriaMetrics
  • Experience building and operating observability platforms at scale
  • Proficiency in Python, Go, or bash for automation and data integration
  • Familiarity with containerized environments and Kubernetes observability
  • Experience with streaming telemetry pipelines (Kafka, OTEL, Promtail, or equivalent)
  • Experience with multi-tenant monitoring architectures
  • Strong written and verbal communication skills
Ideal Experience
  • Experience with GPU observability, particularly NVIDIA DCGM
  • Experience monitoring large-scale GPU or HPC clusters
  • Familiarity with InfiniBand fabric observability
  • Experience building customer-facing or productized infrastructure systems
  • Experience with correlation engines, RCA workflows, or predictive alerting systems
  • Broad exposure to infrastructure domains including networking, storage, and provisioning
ย  Compensation

We are committed to offering competitive compensation that reflects the value each team member brings to our mission. Final offers are based on factors such as experience, skills, geographic location, and role expectations. In addition to base salary, our total rewards package for eligible roles includes a discretionary bonus, a meaningful equity component, and comprehensive benefits.

The anticipated annual base salary range for this role is:$180,000โ€”$200,000 USDBenefits and Perks

We offer a comprehensive and competitive benefits package designed to support our employeesโ€™ health, well-being, and long-term success. Benefits may vary by location, team, and role.

Benefits include:

  • Comprehensive medical, dental and vision coverage (U.S.); Private medical and dental insurance (U.K.)
  • Retirement and financial wellness support (U.S.); Pension contribution (U.K.)
  • Generous paid time off, plus holidays
  • Paid parental leave
  • Professional development support
  • Wellness and work-from-home stipends
  • Flexible work environment

At Lightning AI, we are committed to fostering an inclusive and diverse workplace. We believe that diverse teams drive innovation and create better products. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected characteristic. We are dedicated to building a culture where everyone can thrive and contribute to their fullest potential.