1

Region Reliability Jobs in New York (NOW HIRING)

You will own the design and evolution of major systems within our multi-cloud, multi-region, active ... Own the design, reliability and evolution of core platform applications, mentoring team members on ...

You will own the design and evolution of major systems within our multi-cloud, multi-region, active ... Own the design, reliability and evolution of core platform applications, mentoring team members on ...

next page

Showing results 1-20

Region Reliability information

What is a region reliability engineer?

A Region Reliability Engineer is responsible for ensuring the reliability, availability, and performance of systems or infrastructure within a specific geographic region. They monitor system health, respond to incidents, and implement strategies to minimize downtime and service disruptions. Their work may involve collaborating with local teams, analyzing data to identify trends, and applying best practices in reliability engineering. By proactively addressing potential issues, Region Reliability Engineers help maintain seamless operations and high service standards across their assigned areas.

What are the main challenges region reliability professionals face when coordinating across multiple sites or regions?

Region Reliability professionals often encounter challenges related to aligning maintenance practices and reliability standards across diverse sites with varying operational cultures and equipment. Effective communication and standardization are key, as each region may have distinct processes, equipment types, and reporting structures. Overcoming these challenges typically requires strong collaboration with local teams, proactive sharing of best practices, and leveraging data analytics to identify and address region-specific reliability issues. Building strong relationships and fostering a culture of continuous improvement across all sites are essential for success in this role.

What are the key skills and qualifications needed to thrive as a region reliability engineer, and why are they important?

To thrive as a Region Reliability Engineer, you need strong analytical skills, a background in engineering or a related technical field, and experience with reliability modeling and data analysis. Familiarity with reliability software (like ReliaSoft), maintenance management systems, and relevant certifications such as Certified Reliability Engineer (CRE) are highly valued. Excellent problem-solving abilities, communication, and teamwork are essential soft skills for coordinating with cross-functional teams and driving reliability initiatives. These skills ensure equipment uptime, optimize maintenance strategies, and support safe, efficient operations across multiple sites.

What is the difference between Region Reliability vs Substation Technician?

AspectRegion Reliability
CertificationsTypically requires certifications like NERC Reliability Coordinator or similar
Work EnvironmentPrimarily office-based with some field oversight in utility regions
Employer & IndustryUtilities, power grid management, energy sector
Common Search & ComparisonOften compared for roles in power system reliability and grid management

Region Reliability focuses on overseeing and ensuring the stability of power grids across large regions, often involving planning, monitoring, and compliance. Substation Technicians, on the other hand, perform hands-on maintenance and repairs at specific substations. While both roles are vital in the energy sector, Region Reliability emphasizes system oversight, whereas Substation Technicians focus on equipment and infrastructure maintenance.

What are popular job titles related to Region Reliability jobs in New York?

For Region Reliability jobs in New York, the most frequently searched job titles are:

What cities in New York are hiring for Region Reliability jobs?

Cities in New York with the most Region Reliability job openings:

Infographic showing various Region Reliability job openings in New York as of August 2026, with employment types broken down into 1% As Needed, 53% Full Time, 43% Part Time, and 3% Contract. Highlights an 94% Physical, 1% Hybrid, and 5% Remote job distribution.

Senior Site Reliability Engineer

Jersey City, NJ • On-site

Bank of America
Finance and Insurance • 10K+ employees

$59.50 - $79/hr

Other

PTO

Re-posted 13 days ago


Bank Of America rating

8.3

Company rating: 8.3 out of 10

Based on 538 frontline employees who took The Breakroom Quiz


Job description

Job Overview

At Bank of America, we are guided by a common purpose to help make financial lives better through the power of every connection. We do this by driving Responsible Growth and delivering for our clients, teammates, communities and shareholders every day. Being a Great Place to Work is core to how we drive Responsible Growth. This includes our commitment to being an inclusive workplace, attracting and developing exceptional talent, supporting our teammates’ physical, emotional, and financial wellness, recognizing and rewarding performance, and how we make an impact in the communities we serve.

Key Responsibilities
  • Serve as a senior technical authority for GCP platform reliability, resiliency, observability, automation, and production readiness.
  • Design and implement advanced reliability patterns for Azure landing zones, private networking, DNS, firewalls, regional resiliency, service health checks, and workload onboarding.
  • Lead complex platform reliability initiatives such as secondary-region readiness, egress/ingress observability, private DNS resolver monitoring, GenAI platform health checks, and enterprise dashboard automation.
  • Define and mature SLIs, SLOs, reliability indicators, alerting standards, and service health reporting for Azure platform services.
  • Develop reusable Terraform modules, automation frameworks, and CI/CD patterns that improve consistency, compliance, and operational quality.
  • Drive observability improvements using Log Analytics, Dynatrace, Resource Graph, dashboards, and enterprise monitoring tools.
  • Identify systemic reliability risks and translate them into engineering roadmaps, remediation plans, automation opportunities, and operational controls.
  • Lead deep technical investigations for major incidents, recurring problems, platform defects, or service degradation events.
  • Partner with security and governance teams to integrate IAM, policy-as-code, vulnerability remediation, control validation, and audit readiness into Azure platform operations.
  • Provide technical design input for new Azure services and workloads to ensure operational readiness before production adoption.
  • Mentor SRE engineers and raise the technical bar for automation, troubleshooting, documentation, resiliency design, and production support.
  • Create executive-ready technical summaries, reliability narratives, and recommendations for leadership review.
  • Establish reusable standards for runbooks, dashboards, health checks, problem records, post-incident reviews, and production-readiness gates.
  • Design solutions to visualize key production support metrics enabling operational readiness and site reliability engineering teams to identify scenarios requiring intervention.
  • Develop software solutions and/or improved processes to address work identified as ‘toil’ by collaborating with key partners to identify, track, and remediate processes to free time allocated to reliability.
  • Partner with development and infrastructure teams to create error budget policies prioritizing reliability stories that fall below Service Level Objective thresholds and suggest code optimizations, additional instrumentation, and/or logging structures to gain service reliability visibility.
  • Identify and plan for capacity bottlenecks, vulnerabilities, and opportunities for reliability improvement such as low-level error rates and noise, and reduce manual support effort and/or improve system reliability.
  • Assess monitoring for new changes with development partners and work with monitoring tools team to monitor dashboards and enhance application and system monitoring designs.
  • Engage as a subject matter expert in incident triage efforts, failure scenario modeling, and work with the problem manager to diagnose root causes for complex/high impact incident/problem management investigations.
  • Collaborate with development and infrastructure teams to understand technical solutions and develop Service Level Indicators and SLOs to measure and improve the reliability of the services they support.
Required Qualifications
  • 4+ years of experience in cloud infrastructure engineering, platform engineering, or cloud operations, with exposure to Google Cloud Platform (GCP).
  • Strong hands-on experience with Infrastructure as Code (IaC), including practical use of Terraform or Terraform Enterprise for infrastructure provisioning.
  • Solid understanding of software engineering fundamentals, including version control, code quality, and basic testing practices for infrastructure code.
  • Experience developing and maintaining Terraform modules and infrastructure configurations to support automated cloud environments.
  • Familiarity with CI/CD pipelines for infrastructure deployment, including automated build, test, and release processes.
  • Working knowledge of DevSecOps practices, including integrating security and compliance checks into automated workflows.
  • Good understanding of GCP services and cloud architecture fundamentals, including networking (VPCs, IAM, load balancing).
  • Exposure to policy-as-code, governance, and compliance requirements in enterprise environments.
  • Experience supporting automation and standardization efforts to improve consistency and efficiency in cloud deployments.
  • Understanding of monitoring, logging, and observability tools to support system reliability and performance.
  • Hands‑on experience with incident response, troubleshooting, and root cause analysis in cloud or distributed systems.
Desired Qualifications
  • Ability to collaborate effectively with engineering, architecture, and security teams to support reliable and secure platform operations.
  • Strong problem‑solving and analytical skills, with the ability to diagnose and resolve infrastructure issues.
  • Effective communication skills, with the ability to work within cross‑functional teams and document technical solutions clearly.
  • Interest in learning and applying emerging technologies and automation techniques (including AI/ML where applicable) to improve platform reliability.
Skills
  • Architecture
  • Collaboration
  • Innovative Thinking
  • Result Orientation
  • Solution Design
  • Adaptability
  • Analytical Thinking
  • Influence
  • Stakeholder Management
  • Technical Strategy Development

Shift: 1st shift (United States of America)

Hours per week: 40

Pay range: $152,600.00 – $191,500.00 annualized salary, offers to be determined based on experience, education and skill set. This role is eligible to participate in the annual discretionary plan. Employees are eligible for an annual discretionary award based on their overall individual performance results and behaviors, the performance and contributions of their line of business and/or group, and the overall success of the Company.

Benefits: This role is currently benefits eligible. We provide industry‑leading benefits, access to paid time off, resources and support to our employees so they can make a genuine impact and contribute to the sustainable growth of our business and the communities we serve.

#J-18808-Ljbffr

What Bank Of America employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom


Bank Of America logo

About Bank Of America

Sourced by ZipRecruiter

At Bank of America, we are guided by a common purpose to help make financial lives better through the power of every connection. Responsible Growth is how we run our company and how we deliver for our clients, teammates, communities and shareholders every day. One of the keys to driving Responsible Growth is being a great place to work for our teammates around the world. We're devoted to being a diverse and inclusive workplace for everyone. We hire individuals with a broad range of backgrounds and experiences and invest heavily in our teammates and their families by offering competitive benefits to support their physical, emotional, and financial well-being.

Industry

Finance and insurance

Company size

10,000+ Employees

Headquarters location

Charlotte, NC, US

Social media