1

Service Reliability Engineer Jobs in Raleigh, NC

Site Reliability Engineer II

Raleigh, NC · On-site

$55.50 - $73.75/hr

Site Reliability Engineer II The SRE II sits at the intersection of software engineering and ... Participate in shared on-call rotation covering core platform and infrastructure services; on-call ...

Design, build, and manage our large scale infrastructure and platform services, including public ... Knowledge of SRE principles -- SLOs, error budgets, toil measurement**The Following Are Considered ...

Design, build, and manage our large scale infrastructure and platform services, including public ... Work within a small agile team to develop and improve SRE methodologies, support your peers, plan ...

Senior Site Reliability Engineer

Raleigh, NC · On-site

$55.50 - $73.75/hr

## Senior Site Reliability Engineer### Raleigh-Durham, NCIXL Learning, developer of personalized ... Experience with common service technologies like web servers, message queues, load balancers, and ...

Senior Site Reliability Engineer

Raleigh, NC

$55.50 - $73.75/hr

... Site Reliability Engineer to join our team, and help maintain the reliability and optimal ... Experience with common service technologies like web servers, message queues, load balancers, and ...

Senior Site Reliability Engineer

Raleigh, NC · On-site

$55.50 - $73.75/hr

... Site Reliability Engineer to join our team, and help maintain the reliability and optimal ... Experience with common service technologies like web servers, message queues, load balancers, and ...

Senior Site Reliability Engineer

Raleigh, NC · On-site

$55.50 - $73.75/hr

... Site Reliability Engineer to join our team, and help maintain the reliability and optimal ... Experience with common service technologies like web servers, message queues, load balancers, and ...

We're looking for a Senior SRE to join our Compute Farm team and help build the next generation of our global services platform. At NVIDIA, you'll keep critically important systems running while ...

We're looking for a Senior SRE to join our Compute Farm team and help build the next generation of our global services platform. At NVIDIA, you'll keep critically important systems running while ...

Showing results 21-40

Service Reliability Engineer information

See Raleigh, NC salary details

$54.4K

$105.2K

$125.7K

How much do service reliability engineer jobs pay per year?

As of Sep 14, 2026, the average yearly pay for service reliability engineer in Raleigh, NC is $105,158.00, according to ZipRecruiter salary data. Most workers in this role earn between $91,400.00 and $115,000.00 per year, depending on experience, location, and employer.

What is a service reliability engineer?

Service Reliability Engineers (SREs) are IT professionals who apply software engineering principles to infrastructure and operations problems. Their main goal is to ensure that services are reliable, scalable, and highly available by automating processes, monitoring system performance, and responding to incidents. SREs work closely with development and operations teams to design, build, and maintain robust systems, often using code to manage infrastructure. They also focus on improving system reliability through monitoring, incident response, and post-incident analysis.

How does a service reliability engineer typically collaborate with development and operations teams to improve service uptime?

Service Reliability Engineers (SREs) work closely with both development and operations teams to ensure systems are highly available and resilient. They often participate in incident response, conduct post-incident reviews, and help implement automation to reduce manual intervention. Regular collaboration includes reviewing application changes, contributing to infrastructure design, and sharing best practices for monitoring and alerting. This cross-functional teamwork helps to quickly identify potential issues and proactively enhance system reliability.

What are the key skills and qualifications needed to thrive as a service reliability engineer, and why are they important?

To thrive as a Service Reliability Engineer, you need a solid background in systems administration, networking, coding (often in Python or Go), and experience with cloud infrastructure, typically supported by a degree in computer science or a related field. Familiarity with monitoring tools (like Prometheus), CI/CD pipelines, automation frameworks, and certifications such as AWS Certified DevOps Engineer are highly valued. Strong problem-solving abilities, collaboration, and effective communication skills help you proactively address issues and work well within cross-functional teams. These skills ensure system reliability, quick incident recovery, and the seamless delivery of high-availability services.

What is the difference between Service Reliability Engineer vs Site Reliability Engineer?

AspectService Reliability EngineerSite Reliability Engineer
CredentialsTypically requires experience in software engineering, cloud platforms, and monitoring toolsSimilar credentials, often with a focus on software development and systems engineering
Work EnvironmentWorks closely with development and operations teams to ensure service reliabilityWorks on maintaining and improving system reliability, often in cloud or data center environments
Industry UsageCommon in tech companies focusing on service uptime and customer experienceWidely used in tech, especially in cloud and large-scale infrastructure companies

Both roles focus on ensuring system reliability, often requiring similar skills and certifications. The main difference lies in terminology preference and specific organizational focus, but they generally perform comparable functions in maintaining high service availability.

How much do service reliability engineers get paid?

Service Reliability Engineers typically earn a median annual salary between $90,000 and $130,000, depending on experience, location, and industry. Senior roles or those with specialized skills in cloud platforms and automation tools can command higher compensation, often exceeding $150,000 annually.

What are popular job titles related to Service Reliability Engineer jobs in Raleigh, NC?

For Service Reliability Engineer jobs in Raleigh, NC, the most frequently searched job titles are:

What job categories do people searching Service Reliability Engineer jobs in Raleigh, NC look for?

The top searched job categories for Service Reliability Engineer jobs in Raleigh, NC are:

Infographic showing various Service Reliability Engineer job openings in Raleigh, NC as of September 2026, with employment types broken down into 1% As Needed, 73% Full Time, 20% Part Time, and 6% Contract. Highlights an 84% Physical, 1% Hybrid, and 15% Remote job distribution, with an average salary of $105,158 per year, or $50.6 per hour.

Principal Site Reliability Engineer

Durham, NC • On-site

Fidelity Investments
Investment Management and Consulting Services • 10K+ employees

$55 - $73.25/hr

Full-time

Re-posted 5 days ago


Fidelity Investments rating

8.7

Company rating: 8.7 out of 10

Based on 274 frontline employees who took The Breakroom Quiz


Job description


Note: Fidelity will not provide immigration sponsorship for this position.
Position Description:
Deploys and supports distributed, multi-tiered systems at scale while ensuring high availability and fault tolerance across multiple environments. Builds and operates resilient platforms in Amazon Web Services (AWS) using Elastic Compute Cloud (EC2), Simple Storage Service (S3), and Auto Scaling Groups for dynamic resource management. Designs, develops, and executes performance tests using Java-based frameworks, Apache JMeter, k6, and Rush-hour to validate system behavior under day-to-day traffic patterns. Defines and implements observability practices to monitor system health, latency, and error rates through metrics, logs, and distributed tracing using Datadog, Grafana, Splunk, and the Elasticsearch, Logstash, and Kibana (ELK) stack. Automates operational workflows with Python and Shell scripting to enhance efficiency and reduce manual tasks. Supports consistent build, deployment, and orchestration processes using cloud computing and DevOps technologies -- Continuous Integration and Continuous Delivery (CI/CD) pipelines and Kubernetes. Supports Site Reliability Engineering (SRE) functions by establishing Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets, and implementing proactive monitoring and incident response strategies. Builds and refines methodologies for performance, load, stress, and chaos testing and develops analytics and reports aligned with business needs to improve system resilience and optimization.
Primary Responsibilities:
  • Defines and leads enterprise-level reliability strategies.
  • Architects resilient systems and infrastructure.
  • Creates and publishes performance test results report with recommendations on quality improvement.
  • Maintains scalability and resiliency of complex environment.
  • Implements advanced observability practices and techniques at scale.
  • Manages and interprets large datasets using query languages and visualization tools.
  • Advises senior leadership on reliability engineering best practices.
  • Mentors junior engineers.
  • Performs independent and complex technical and functional analysis for multiple divisional initiatives.
  • Develops innovative solutions to improve system availability, scalability, and performance.
  • Designs, implements, and maintains performance test frameworks.

Education and Experience:
Bachelor's degree in Computer Science, Engineering, Information Technology Management, Information Systems Security, Business Administration, or a closely related field (or foreign education equivalent) and five (5) years of experience as a Principal Site Reliability Engineer (or closely related occupation) implementing highly available trading systems in a financial services environment.
Or, alternatively, Master's degree in Computer Science, Engineering, Information Technology Management, Information Systems Security, Business Administration, or a closely related field (or foreign education equivalent) and three (3) years of experience as a Principal Site Reliability Engineer (or closely related occupation) implementing highly available trading systems in a financial services environment.
Skills and Knowledge:
Candidate must also possess:
  • Demonstrated Expertise ("DE") performing software performance benchmarking and engineering for online financial web applications, Application Programming Interfaces (APIs), and mobile transactions according to DevOps practices, using performance benchmarking tools Rushhour, Locust, K6, and JMeter; and configuring CI/CD and test automation, using Jenkins, Sonar, Ant, Maven, Artifactory, and Terraform in AWS.
  • DE solutioning, designing, architecting, and building scalable and resilient enterprise-grade software platforms using cloud-based architecture and AWS services (EC2, Elastic Container Service (ECS), Lambda, Elastic MapReduce (EMR), and CloudFormation); developing microservices on Elastic Kubernetes Service (EKS), implementing CI/CD pipelines using DevOps tools (Bitbucket, GitHub, Artifactory, Sonar, Veracode, and Helm), and adhering to DevOps practices along with leveraging Java, Python, Spring Boot, Docker, EKS, and AWS.
  • DE analyzing and monitoring system and application performance across Apache, NGINX, Java, and Node.js platforms, and Linux and Windows environments, using Splunk, Datadog, Kibana, Grafana, and AWS CloudWatch; diagnosing performance bottlenecks, recommending tuning strategies, reducing Mean Time to Detect (MTTD) and Mean Time to Repair (MTTR), using Application Performance Monitoring (APM) tools -- Dynatrace, New Relic, Splunk, and Datadog; and performing capacity planning to optimize Central Processing Unit (CPU), memory, and process configurations.
  • DE instrumenting advanced observability practices at scale across cloud-native and hybrid environments; defining and tracking SLOs and SLIs to ensure reliability and performance metrics, using Python automation, Infrastructure as Code (IaC) methodologies, and observability tools (Datadog, Splunk, Dynatrace, Grafana, and the ELK stack; developing custom dashboards, alerting rules, and automated incident response workflows to proactively detect and resolve performance degradations, using Datadog, Catchpoint, Grafana, ELK stack, and Cloudwatch; and enabling actionable insights through trace-level correlation of end-to-end (E2E) user journeys and system behaviors, using Dynatrace, Splunk, Draw.io, and Miro.

#PE1M2
#LI-DNI
Fidelity's Onsite Working Model
Fidelity is transitioning to a full-time onsite working model through a phased rollout across regions and roles. Currently, some roles and locations require 100% onsite presence, while others require less. Onsite expectations are likely to evolve as the rollout continues. This transition does not apply to fully remote roles.
Certifications:
Category:
Information Technology
Please be advised that Fidelity's business is governed by the provisions of the Securities Exchange Act of 1934, the Investment Advisers Act of 1940, the Investment Company Act of 1940, ERISA, numerous state laws governing securities, investment and retirement-related financial activities and the rules and regulations of numerous self-regulatory organizations, including FINRA, among others. Those laws and regulations may restrict Fidelity from hiring and/or associating with individuals with certain Criminal Histories.

What Fidelity Investments employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom