1

Observability Sre Jobs in Raleigh, NC (NOW HIRING)

Senior SRE Engineer

Raleigh, NC · On-site

$55.50 - $73.75/hr

As a Senior SRE Engineer, you will: * Design, implement, and maintain scalable systems for uptime ... Observability Platforms: Hands-on experience with Datadog for metrics, distributed tracing, log ...

Site Reliability Engineer

Morrisville, NC · On-site

$120K - $150K/yr

Description Site Reliability Engineer The Company: Varonis (Nasdaq: VRNS) secures AI and the data that powers it. The Varonis platform gives organizations automated visibility and control over their ...

... observability or data-driven operations (AIOps/ML-driven signals) that materially reduce manual ... HPC/SRE problems and their solutions. * Maintainer or comaintainer responsibilities for an open ...

Senior Site Reliability Engineer

Raleigh, NC · On-site

$55.50 - $73.75/hr

... Site Reliability Engineer to join our team, and help maintain the reliability and optimal ... Fluency in observability tools and methodologies * Flexibility to work with a variety of ...

Senior Site Reliability Engineer

Raleigh, NC · On-site

$55.50 - $73.75/hr

... Site Reliability Engineer to join our team, and help maintain the reliability and optimal ... Fluency in observability tools and methodologies * Flexibility to work with a variety of ...

Senior SRE Engineer

Raleigh, NC · On-site

$55.50 - $73.75/hr

As a Senior SRE Engineer, you will: * Design, implement, and maintain scalable systems for uptime ... Observability Platforms: Hands-on experience with Datadog for metrics, distributed tracing, log ...

The Role We are seeking a driven and development-focused Site Reliability Engineer to join our SRE department. This group ensures that our software applications and infrastructure are reliable ...

Site Reliability Engineering Lead

Raleigh, NC

$55.50 - $73.75/hr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

The role involves standardizing observability practices, mentoring SRE team members, and contributing to enterprise-wide reliability frameworks. Candidates require 7+ years of experience, expertise ...

Senior Site Reliability Engineer

Raleigh, NC · On-site

$118K - $195K/yr

  • Medical

  • Dental

  • Vision

  • Retirement

  • PTO

Job Summary The Red Hat IT OpenShift team is looking for a Senior Site Reliability Engineer (SRE) to design, develop, scale, and operate our Red Hat Hybrid OpenShift Platforms (on-prem & cloud). As a ...

Site Reliability Engineer The Company:Varonis(Nasdaq: VRNS) secures AI and the data that powers it. The Varonis platform gives organizations automated visibility and control over their critical data ...

Site Reliability Engineer Intern 2027

Durham, NC · On-site

$55 - $73.25/hr

Your primary responsibilities include: • 24x7 Observability: Be part of a worldwide team that ... Job Title Site Reliability Engineer Intern 2027 Date posted 11-Aug-2026 Job ID 128513 City ...

New

Showing results 21-40

Observability Sre information

See Raleigh, NC salary details

$10

$61

$89

How much do observability sre jobs pay per hour?

As of Aug 16, 2026, the average hourly pay for observability sre in Raleigh, NC is $61.96, according to ZipRecruiter salary data. Most workers in this role earn between $53.27 and $70.82 per hour, depending on experience, location, and employer.

What is an Observability SRE?

An Observability SRE (Site Reliability Engineer) is a specialist focused on ensuring that systems and applications are transparent, measurable, and reliable. Their main responsibility is to implement and maintain tools for monitoring, logging, and tracing, providing insights into system performance and health. Observability SREs help teams quickly detect, diagnose, and resolve issues by making system behavior visible and understandable. They play a critical role in uptime, incident response, and performance optimization, bridging the gap between software development and IT operations.

What are the key skills and qualifications needed to thrive as an Observability SRE?

To thrive as an Observability SRE, you need a solid background in systems engineering, monitoring best practices, and expertise in observability concepts, often supported by a degree in computer science or related fields. Familiarity with tools like Prometheus, Grafana, ELK stack, and cloud monitoring platforms, as well as scripting languages such as Python or Bash, is typically required. Strong problem-solving, collaboration, and communication skills help SREs respond to incidents and work across teams effectively. These skills ensure system reliability, rapid issue detection, and continuous service improvement in complex technical environments.

What are some typical challenges faced by Observability SREs when implementing monitoring solutions across diverse systems?

Observability SREs often encounter challenges when integrating monitoring tools across varied technology stacks and legacy systems. Ensuring consistent data collection, standardizing metrics, and maintaining visibility in complex, distributed environments can be difficult. Collaborating with development and operations teams to define meaningful alerts and dashboards requires strong communication and a deep understanding of both infrastructure and application behaviors. Staying up-to-date with evolving tools and best practices is also essential to address emerging observability needs.

What is the difference between Observability Sre vs Site Reliability Engineer?

AspectObservability SreSite Reliability Engineer
Primary FocusMonitoring, logging, and tracing to ensure system observabilitySystem reliability, automation, and infrastructure management
Skills & CertificationsMonitoring tools, scripting, cloud platforms, observability frameworksLinux, scripting, cloud services, automation tools
Work EnvironmentCollaborates with SRE, DevOps, and development teams on observability practicesBuilds and maintains scalable, reliable systems in production

While both roles focus on system stability, Observability Sre specializes in monitoring and diagnostics, whereas Site Reliability Engineers focus on overall system reliability and automation. They often work together to ensure robust, observable, and reliable systems.

What are popular job titles related to Observability Sre jobs in Raleigh, NC?

For Observability Sre jobs in Raleigh, NC, the most frequently searched job titles are:

What job categories do people searching Observability Sre jobs in Raleigh, NC look for?

The top searched job categories for Observability Sre jobs in Raleigh, NC are:

Infographic showing various Observability Sre job openings in Raleigh, NC as of August 2026, with employment types broken down into 65% Full Time, and 35% Contract. Highlights an 68% In-person, and 32% Remote job distribution, with an average salary of $128,882 per year, or $62 per hour.

Site Reliability Engineering Lead

Habitat For Humanity Of Durham

Raleigh, NC • On-site

$120 - $180/hr

Other

Medical, Dental, Vision, Life, Retirement, PTO

Posted yesterday

New


Job description

If you have a disability and need assistance with the application, you can request a reasonable accommodation. Send an email to Accessibility (accommodation requests only; other inquiries won't receive a response).

Regular or Temporary: Regular

Language Fluency: English (Required)

Work Shift: 1st shift (United States of America)

Please review the following job description:

The Site Reliability Engineering Lead role focuses on enhancing the reliability and operational excellence of enterprise platforms across hybrid cloud and on-premises environments. This senior technical leader drives improvements in automation, observability, and incident management while collaborating across multiple business and technology teams.

Responsibilities include leading major incident responses, driving problem management, and implementing automation to reduce service downtime.

The role involves standardizing observability practices, mentoring SRE team members, and contributing to enterprise-wide reliability frameworks.

Candidates require 7+ years of experience, expertise in distributed systems, Kubernetes, automation scripting, and strong leadership in incident management.

ESSENTIAL DUTIES AND RESPONSIBILITIES
  • Implements software architecture and engineering approaches for complex initiatives within the job area, contributing to technical plans and working to achieve operational targets with major impact on results.
  • Adopts and refines advanced software engineering standards, practices, and governance mechanisms for the job area, influencing how multiple teams improve quality, reliability, and delivery.
  • Collaborates with senior engineers, product partners, and architecture teammates to shape technology approaches for the domain, providing deep technical insight and proposing solution patterns that inform local roadmaps and priorities.
  • Leads the end-to-end technical design and implementation of scalable, secure, and highly available software solutions for the job area, producing patterns and examples that other technical professionals can follow.
  • Independently troubleshoots and resolves complex technical issues in the area of responsibility, designing innovative architectures and performance, reliability, and scalability improvements that advance business objectives.
  • Provides ongoing technical guidance, coaching, and training to other engineers, delegating and reviewing work from lower-level technical professionals and raising the technical bar through design reviews and knowledge sharing.
  • Evaluates emerging technologies and techniques relevant to the job area, building prototypes and solution concepts that contribute measurable input into new features, products, or capabilities.
  • Contributes to the development of long-term technical goals and plans for the area of responsibility through well-reasoned recommendations, design proposals, and implementation experience.
  • Leads large or complex initiatives within the job area, coordinating and delegating technical work that may span outside the immediate team, and ensuring cohesive, high-quality outcomes with limited supervision.
Qualifications Required Qualifications
  • Bachelor's degree in Computer Science, Software Engineering, or related field.
  • Minimum of 7 years of professional experience in software development.
  • Deep knowledge of multiple programming languages, software architecture, and design principles.
  • Deep understanding of software development lifecycle, testing, deployment, and security practices.
Preferred Qualifications
  • Advanced degree in Computer Science or related technical discipline.
  • Professional certifications such as Certified Software Development Professional (CSDP) or equivalent.
  • Deep expertise in cloud-native architectures, microservices, container orchestration, and DevOps.
  • Strong familiarity with Agile frameworks, continuous integration/continuous deployment (CI/CD), and enterprise innovation management.
  • 7+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Infrastructure Operations.
  • Deep hands-on experience with distributed systems, container orchestration (Kubernetes), and cloud-native operational tooling.
  • Proficiency with automation and scripting languages (Python, Go, PowerShell, Ansible).
  • Strong understanding of observability platforms (Splunk, Dynatrace) and event-driven monitoring.
  • Proven leadership in major incident management and cross-team technical coordination.
  • Strong grasp of networking, Linux/Unix internals, and modern infrastructure patterns.
  • Excellent communication skills, including executive-level situational awareness during critical incidents.
  • Demonstrated ability to influence technical roadmaps and drive adoption of reliability best practices.
  • Financial services or regulated industry experience.
  • Experience enabling large-scale SRE transformations or modernization initiatives.
  • Familiarity with chaos engineering, resilience assessments, and service failure modeling.
  • Exposure to hybrid-cloud and multi-cloud operational frameworks.
  • Experience contributing to or leading Center for Enablement functions or Communities of Practice.
Key Responsibilities Incident & Problem Management Leadership
  • Lead major and high-severity incident response efforts, focusing on diagnosing technical root causes therein, and driving multi-team technical resolution.
  • Drive problem management to closure, ensuring systemic fixes replace recurring operational risks.
  • Establish and maintain standardized incident playbooks, escalation paths, and communication frameworks.
Reliability Engineering & Automation
  • Architect and deliver automation solutions that eliminate toil, reduce MTTR, and increase service resilience.
  • Implement intelligent alerting, anomaly detection, and event correlation leveraging AI and AIOps tools.
  • Guide and enforce SLO/SLI adoption across product teams, ensuring metrics inform decision-making and prioritization.
Observability & Operational Excellence
  • Enhance telemetry coverage across logs, metrics, traces, and events using platforms such as Dynatrace and Splunk.
  • Define and standardize enterprise observability practices, dashboards, and KPIs.
  • Ensure operational readiness of applications and platforms through resiliency testing, chaos engineering, and failure-mode validation.
Cross-Functional Leadership & Influence
  • Partner with Delivery, Architecture, Security, and Risk teams to embed reliability and resilience into design and execution.
  • Act as a change agent to elevate operational maturity and drive transformative improvements across Wholesale.
  • Lead workshops, maturity assessments, and enablement sessions through the SRE C4E and Communities of Practice.
Standardization & Documentation
  • Develop, maintain, and enforce runbooks, response playbooks, and automated recovery patterns.
  • Contribute to enterprise SRE frameworks, templates, and maturity models.
  • Promote consistent adoption of best practices across domains and lines of business.
Mentorship & Technical Development
  • Coach and mentor Associate, Professional, and Senior SREs to build technical depth and operational discipline.
  • Provide thought leadership in SRE methodologies, cloud-native operational patterns, and automated reliability engineering.

For this opportunity, Truist will not sponsor an applicant for work visa status or employment authorization, nor will we offer any immigration-related support for this position (including, but not limited to H-1B, F-1 OPT, F-1 STEM OPT, F-1 CPT, J-1, TN-1 or TN-2, E-3, O-1, or future sponsorship for U.S. lawful permanent residence status.)

Candidate must be willing to work onsite Monday - Friday at either office in Charlotte NC, Raleigh NC, or Atlanta, GA.

General Description of Available Benefits for Eligible Employees of Truist Financial Corporation: All regular teammates (not temporary or contingent workers) working 20 hours or more per week are eligible for benefits, though eligibility for specific benefits may be determined by the division of Truist offering the position. Truist offers medical, dental, vision, life insurance, disability, accidental death and dismemberment, tax-preferred savings accounts, and a 401k plan to teammates. Teammates also receive no less than 10 days of vacation (prorated based on date of hire and by full-time or part-time status) during their first year of employment, along with 10 sick days (also prorated), and paid holidays. For more details on Truists generous benefit plans, please visit our Benefits site. Depending on the position and division, this job may also be eligible for Truists defined benefit pension plan, restricted stock units, and/or a deferred compensation plan. As you advance through the hiring process, you will also learn more about the specific benefits available for any non-temporary position for which you apply, based on full-time or part-time status, position, and division of work.

Truist is an Equal Opportunity Employer that does not discriminate on the basis of race, gender, color, religion, citizenship or national origin, age, sexual orientation, gender identity, disability, veteran status, or other classification protected by law. Truist is a Drug Free Workplace.

EEO is the Law E-Verify IER Right to Work

#J-18808-Ljbffr