1

Director Observability Jobs in Portland, OR (NOW HIRING)

Account Executive - Splunk

Portland, OR · On-site +1

$212K - $291K/yr

Leading enterprises use our unified security and observability platform to keep their digital ... Minimum Qualifications * 5+ years of direct sales experience selling enterprise software to large ...

Account Executive - Splunk

Portland, OR · On-site +1

$212K - $291K/yr

Leading enterprises use our unified security and observability platform to keep their digital ... Minimum Qualifications * 5+ years of direct sales experience selling enterprise software to large ...

Showing results 21-27

Director Observability information

What does a director of observability do?

A Director of Observability leads the strategy and implementation of monitoring, logging, and tracing systems to ensure the health and performance of technical infrastructure. They work with engineering and operations teams to develop best practices, select appropriate tools, and set standards for observability across the organization. Their goal is to provide visibility into system behavior, quickly identify and resolve incidents, and support continuous improvement in system reliability and performance.

How does a director of observability typically collaborate with engineering and operations teams to drive organizational goals?

A Director of Observability works closely with engineering and operations teams to ensure that systems are monitored effectively and issues are identified and resolved quickly. This collaboration often involves developing unified monitoring strategies, aligning observability tools and processes, and facilitating incident response post-mortems. The Director also leads cross-functional meetings to establish best practices, set key performance indicators (KPIs), and ensure observability is integrated into the software development lifecycle. By acting as a bridge between technical teams, they help foster a culture of transparency, reliability, and continuous improvement.

What are the key skills and qualifications needed to thrive as a director of observability, and why are they important?

To thrive as a Director of Observability, you need deep expertise in monitoring, logging, and distributed systems, typically backed by a degree in computer science or a related field and extensive experience in IT or DevOps leadership roles. Proficiency with observability tools such as Prometheus, Grafana, Datadog, Splunk, and APM solutions, along with knowledge of cloud platforms and relevant certifications, is essential. Strong leadership, strategic thinking, and communication skills help drive cross-functional initiatives and foster a culture of reliability. These skills and qualities are crucial for ensuring system health, rapid incident response, and alignment between technical teams and organizational objectives.

What is the difference between Director Observability vs Site Reliability Engineer?

AspectDirector ObservabilitySite Reliability Engineer
Primary FocusOversees observability strategies, tools, and teams to ensure system visibility and performanceBuilds and maintains reliable systems, automates deployment, and manages incident response
CredentialsTypically requires advanced knowledge of monitoring, cloud platforms, and leadership experienceOften has software engineering background, with skills in scripting, automation, and systems engineering
Work EnvironmentLeads teams in tech companies, focusing on monitoring and analytics toolsWorks closely with development and operations teams to ensure system reliability

While both roles focus on system performance and reliability, the Director Observability primarily manages observability strategies and teams, whereas the Site Reliability Engineer is hands-on, building and maintaining reliable systems. The roles complement each other in ensuring optimal system performance and uptime.

What are popular job titles related to Director Observability jobs in Portland, OR?

For Director Observability jobs in Portland, OR, the most frequently searched job titles are:

What job categories do people searching Director Observability jobs in Portland, OR look for?

The top searched job categories for Director Observability jobs in Portland, OR are:

What cities near Portland, OR are hiring for Director Observability jobs?

Cities near Portland, OR with the most Director Observability job openings:

Senior Development Operations Engineer (DevOps) (Remote)

Rapta, Inc

Lake Oswego, OR • Remote

$70K - $90K/yr

Full-time

Posted 26 days ago


Job description

Senior Development Operations Engineer (DevOps)

Full-Time Position | Remote (US)

About Us

Rapta is revolutionizing American manufacturing with our AI-powered vision systems. Our cutting-edge software platform expands manufacturing capacity 30% by reducing errors 90%+ and automating quality control and inspection processes. We're looking for a talented senior development operations engineer to join our mission and make a real impact.

Position Overview

We're seeking an experienced senior development operations (DevOps) engineer to own the build, test, and deployment pipeline for our distributed computer vision platform running at edge sites across customer manufacturing floors. This is a full-time remote position working directly with our engineering team to harden our release process, drive deployment automation, and pioneer LLM-driven test generation and validation at Rapta.

What You'll Do

  • Own and evolve the end-to-end release pipeline branching strategy, build orchestration, artifact promotion, and rollback across our Bazel monorepo and Python deployable units
  • Design and maintain Ansible-driven fleet automation for heterogeneous Linux edge nodes (Ubuntu LTS, NVIDIA driver stacks, Docker with NVIDIA runtime)
  • Manage all update tooling, currently written in Golang
  • Build LLM-powered automated testing systems: test generation from specs, flake triage, log/failure analysis, regression diffing, and release-note synthesis from commit and ticket history
  • Harden CI/CD for offline and bandwidth-constrained deployment targets (airgap wheel distribution, signed artifacts, deterministic builds)
  • Drive observability for releases deployment telemetry, version drift detection, and post-deploy health validation across the fleet
  • Mentor engineers on release hygiene, reproducible builds, and infrastructure-as-code practices

What We're Looking For

  • 10+ years of professional experience in release engineering, DevOps, or SRE roles shipping production Linux systems
  • Deep curiosity for software, infrastructure, and applied AI particularly using LLMs as production engineering tools, not just chat assistants
  • Expert-level Python (3.8+) with a strong grasp of packaging, dependency resolution, and PEP 440 versioning discipline
  • Demonstrated ownership of Linux fleets at scale kernel, systemd, networking, package management
  • Excellence in technical communication, runbook authorship, and post-incident documentation
  • Strong systems thinking comfortable reasoning about failure modes across hardware, OS, container, and application layers

Required Technical Skills

  • Expert proficiency with Ansible (roles, dynamic inventory, idempotent design); working knowledge of Terraform
  • Expert proficiency with Docker, including creation and lifecycle management of containers, image hardening, registry management and installing & configuring the NVIDIA container runtime
  • Production experience with Linux administration: systemd, networking (VLANs, DHCP, DNS), kernel/driver management (especially NVIDIA/DKMS), package and APT internals
  • Strong Python skills focused on tooling, automation, packaging (wheels, pip, private indexes), and subprocess/CI integration
  • Proficiency with Git workflows, branching strategies, and modern CI/CD systems (GitHub Actions, GitLab CI, or equivalent)
  • Experience designing and operating automated test infrastructure unit, integration, hardware-in-the-loop, and end-to-end
  • Practical experience using LLMs (Anthropic, OpenAI, or local) as part of engineering workflows test generation, code review augmentation, log analysis, or agentic tooling

Nice to Have

  • Bazel or similar monorepo build systems
  • Edge or embedded deployment experience
  • Tailscale, WireGuard, or zero-trust networking in production
  • gRPC/protobuf service ecosystems
  • Vault, PKI, or secrets management at fleet scale
  • Background in regulated or compliance-driven environments (CMMC, ISO 27001, SOC 2)

Why Join Us?

  • Work on cutting-edge AI infrastructure with real-world impact on American manufacturing
  • Build the release engineering foundation for a late-seed company actively scaling
  • Direct collaboration with the CTO and engineering leadership
  • Remote-first culture with flexible hours
  • Meaningful equity in a company solving hard problems

Location Requirements

  • Remote (US)

Equal Opportunity

Rapta is committed to hiring and retaining a diverse workforce. We are proud to be an Equal Opportunity/Affirmative Action Employer, making decisions without regard to race, color, religion, creed, sex, sexual orientation, gender identity, marital status, national origin, age, veteran status, disability, or any other protected class.

How to Apply No recruiters or agencies we only accept applications directly from applicants.