1

Service Reliability Engineer Jobs in New York (NOW HIRING)

Showing results 41-60

Service Reliability Engineer information

See New York salary details

$66.7K

$129.1K

$154.3K

How much do service reliability engineer jobs pay per year?

As of Aug 17, 2026, the average yearly pay for service reliability engineer in New York is $129,066.00, according to ZipRecruiter salary data. Most workers in this role earn between $112,100.00 and $141,100.00 per year, depending on experience, location, and employer.

What is a service reliability engineer?

Service Reliability Engineers (SREs) are IT professionals who apply software engineering principles to infrastructure and operations problems. Their main goal is to ensure that services are reliable, scalable, and highly available by automating processes, monitoring system performance, and responding to incidents. SREs work closely with development and operations teams to design, build, and maintain robust systems, often using code to manage infrastructure. They also focus on improving system reliability through monitoring, incident response, and post-incident analysis.

How does a service reliability engineer typically collaborate with development and operations teams to improve service uptime?

Service Reliability Engineers (SREs) work closely with both development and operations teams to ensure systems are highly available and resilient. They often participate in incident response, conduct post-incident reviews, and help implement automation to reduce manual intervention. Regular collaboration includes reviewing application changes, contributing to infrastructure design, and sharing best practices for monitoring and alerting. This cross-functional teamwork helps to quickly identify potential issues and proactively enhance system reliability.

What are the key skills and qualifications needed to thrive as a service reliability engineer, and why are they important?

To thrive as a Service Reliability Engineer, you need a solid background in systems administration, networking, coding (often in Python or Go), and experience with cloud infrastructure, typically supported by a degree in computer science or a related field. Familiarity with monitoring tools (like Prometheus), CI/CD pipelines, automation frameworks, and certifications such as AWS Certified DevOps Engineer are highly valued. Strong problem-solving abilities, collaboration, and effective communication skills help you proactively address issues and work well within cross-functional teams. These skills ensure system reliability, quick incident recovery, and the seamless delivery of high-availability services.

What is the difference between Service Reliability Engineer vs Site Reliability Engineer?

AspectService Reliability EngineerSite Reliability Engineer
CredentialsTypically requires experience in software engineering, cloud platforms, and monitoring toolsSimilar credentials, often with a focus on software development and systems engineering
Work EnvironmentWorks closely with development and operations teams to ensure service reliabilityWorks on maintaining and improving system reliability, often in cloud or data center environments
Industry UsageCommon in tech companies focusing on service uptime and customer experienceWidely used in tech, especially in cloud and large-scale infrastructure companies

Both roles focus on ensuring system reliability, often requiring similar skills and certifications. The main difference lies in terminology preference and specific organizational focus, but they generally perform comparable functions in maintaining high service availability.

What are popular job titles related to Service Reliability Engineer jobs in New York?

For Service Reliability Engineer jobs in New York, the most frequently searched job titles are:

What job categories do people searching Service Reliability Engineer jobs in New York look for?

The top searched job categories for Service Reliability Engineer jobs in New York are:

Infographic showing various Service Reliability Engineer job openings in New York as of August 2026, with employment types broken down into 1% As Needed, 76% Full Time, 20% Part Time, 1% Temporary, and 2% Contract. Highlights an 96% Physical, 1% Hybrid, and 3% Remote job distribution, with an average salary of $129,066 per year, or $62.1 per hour.

Site Reliability Engineer (SRE)

Long Finch Technologies

South Richmond Hill, NY โ€ข On-site

$59.75 - $79.50/hr

Full-time

This job post hasย expired today.ย Applications are no longer accepted.


Job description

Overview

We are seeking an experienced Site Reliability Engineer (SRE) โ€“ Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise with modern Site Reliability Engineering practices to improve platform reliability, scalability, performance, automation, and operational excellence.

The ideal candidate will be responsible for ensuring platform reliability through infrastructure automation, OS upgrades, proactive monitoring, incident response, capacity planning, and continuous service improvement while collaborating with infrastructure, security, and application teams.

Key Responsibilities

  • Operate enterprise-scale private cloud infrastructure built on Microsoft Hyper-V.
  • Optimize, and support highly available VDI environments on Hyper-V.
  • Improve platform reliability, availability, scalability, and resiliency by applying SRE principles and engineering best practices.
  • Disaster recovery, backup, patch management, and business continuity strategies.
  • Define and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational metrics for critical infrastructure services.
  • Automate infrastructure provisioning, configuration management, and operational workflows using PowerShell and Infrastructure as Code (IaC) principles wherever applicable.
  • Manage Hyper-V Failover Clusters, host lifecycle, storage, networking, and capacity to ensure high availability and business continuity.
  • Develop proactive monitoring, alerting, logging, and observability capabilities to detect and prevent service degradation.
  • Lead incident response for infrastructure-related outages, perform root cause analysis (RCA), and implement preventive actions through post-incident reviews.
  • Perform capacity planning, performance tuning, and resource optimization across Hyper-V clusters and VDI platforms.
  • Support infrastructure migration initiatives including P2V, V2V, workload modernization, and private cloud transformations.
  • Collaborate closely with Security, Networking, Platform Engineering, and Application teams to improve platform reliability and operational efficiency.
  • Develop and maintain technical documentation, architecture diagrams, operational runbooks, automation scripts, and standard operating procedures.
  • Mentor junior engineers and promote SRE culture, automation, and operational best practices across the team.


Experience & Qualifications

  • 6+ years of infrastructure engineering experience with at least 4+ years of hands-on Microsoft Hyper-V administration.
  • Demonstrated experience operating mission-critical enterprise infrastructure with high availability and reliability requirements.
  • Proven experience implementing automation to reduce operational overhead and improve service reliability.
  • Experience supporting enterprise private cloud and VDI environments.
  • Experience participating in incident response, root cause analysis, and continuous service improvement initiatives.
  • Microsoft certifications such as Microsoft Certified: Windows Server Hybrid Administrator Associate or equivalent are desirable.
  • Experience in Banking or Financial Services environments is advantageous.

    Preferred Skills

    • Windows Server 2016/2019/2022 administration.
    • Experience with backup and disaster recovery solutions such as Veeam, Altaro, or native Hyper-V Replica.
    • Exposure to hybrid cloud and private cloud platforms.
    • Familiarity with monitoring and observability platforms such as SCOM, Azure Monitor, Prometheus, Grafana, Splunk, or similar tools.
    • Experience supporting enterprise VDI environments.
    • Understanding of ITIL Incident, Problem, Change, and Release Management.
    • Experience working in regulated industries such as Banking or Financial Services.