1

Azure Site Reliability Engineer Jobs in Virginia

Position: Site Reliability Engineer IV Experience: 9-12 Years Location: Bangalore (Ecospace) Candescent Site Reliability Engineering (SRE) mission is to proactively ensure the reliability ...

New

SRE Engineer

Arlington, VA · On-site

$65.75 - $87.25/hr

Spatial Front, Inc. is a recognized workplace seeking a SRE Engineer to enhance their Infrastructure, Production, and Compliance Support team. The role focuses on improving the reliability and ...

Site Reliability Engineer - Networking

Ashburn, VA · On-site

$58.25 - $77.50/hr

As a Site Reliability Engineer, you will be focused on supporting a specific, highly available, and ... AWS/Azure, Docker, K8s, Ansible, REST APIs, TLS or Cloud/ISP/Telco exposure. Preferred ...

Site Reliability Engineer - Networking

Herndon, VA · On-site

$58.50 - $78/hr

As a Site Reliability Engineer, you will be focused on supporting a specific, highly available, and ... AWS/Azure, Docker, K8s, Ansible, REST APIs, TLS or Cloud/ISP/Telco exposure. Preferred ...

Site Reliability Engineer - Networking

Arlington, VA · On-site

$65.50 - $87.25/hr

As a Site Reliability Engineer, you will be focused on supporting a specific, highly available, and ... AWS/Azure, Docker, K8s, Ansible, REST APIs, TLS or Cloud/ISP/Telco exposure. Preferred ...

Site Reliability Engineer

Chantilly, VA · On-site

$62K - $141K/yr

Site Reliability Engineer The Opportunity: Engineering to make a system more resilient and efficient frees up time and money to build more capabilities. Whether you come from a background in network ...

Showing results 21-40

Azure Site Reliability Engineer information

What is the difference between Azure Site Reliability Engineer vs Cloud Engineer?

AspectAzure Site Reliability EngineerCloud Engineer
CertificationsAzure certifications, SRE-related skillsCloud platform certifications (Azure, AWS, GCP)
Work EnvironmentFocus on reliability, monitoring, automation in AzureDesign, implement, manage cloud infrastructure across platforms
Industry UsageTech companies using Azure for scalable, reliable servicesOrganizations adopting cloud solutions, multi-cloud environments

Azure Site Reliability Engineers specialize in maintaining and improving the reliability of Azure-based systems, focusing on automation and monitoring. Cloud Engineers have a broader scope, working across various cloud platforms to design and manage cloud infrastructure. While both roles require cloud certifications and involve cloud environments, SREs emphasize system reliability within Azure, whereas Cloud Engineers focus on overall cloud architecture and deployment.

What job categories do people searching Azure Site Reliability Engineer jobs in Virginia look for? The top searched job categories for Azure Site Reliability Engineer jobs in Virginia are:
Infographic showing various Azure Site Reliability Engineer job openings in Virginia as of August 2026, with employment types broken down into 33% Full Time, and 67% Contract. Highlights an 100% In-person job distribution.

Site Reliability Engineer IV

Capitolis

Sterling, VA • On-site

$120 - $160/hr

Other

Posted 3 days ago

New


Job description

Candescent is the leading cloud-based digital banking solutions provider for financial institutions. We are transforming digital banking with intelligent, cloud-powered solutions that connect account opening, digital banking, and branch experiences for financial institutions. Our advanced technology and developer tools enable seamless, differentiated customer journeys that elevate trust, service, and innovation. Success here requires flexibility in a fast-paced environment, a client-first mindset, and a commitment to delivering consistent, reliable results as part of a performance-driven, values-led team. With team members around the world, Candescent is an equal opportunity employer.

Position: Site Reliability Engineer IV

Experience: 9-12 Years

Location: Bangalore (Ecospace)

Candescent Site Reliability Engineering (SRE) mission is to proactively ensure the reliability, availability and performance of our Digital First banking applications. As a member of the SRE team, you will focus on building and operating highly reliable application platforms by applying SRE principles such as automation, observability, resilience and continuous improvement.

You will partner closely with application and platform teams to define reliability standards, implement monitoring, alerting and incident response practices and embed scalability and performance considerations into application design and delivery. Through tooling, automation, and best practices, you will help development teams build and operate services that meet agreed reliability objectives.

As a senior engineer in the organization, you will also provide mentorship within the SRE team and across peer engineering teams, helping elevate operational maturity, drive adoption of SRE practices, and strengthen reliability culture across our core initiatives.

Responsibilities
  • Support and operate production applications running on Kubernetes and AWS
  • Troubleshoot application-level issues using logs, metrics, traces, and runtime signals
  • Participate in incident response, root cause analysis, and post-incident reviews
  • Work closely with development teams to understand application architecture, dependencies, and data flows
  • Improve application observability by defining meaningful alerts, dashboards, and SLOs
  • Automate repetitive operational tasks to reduce toil
  • Support application deployments, rollbacks, and runtime configuration changes
  • Identify reliability, performance, and scalability gaps in application behavior
  • Drive continuous improvements in operational readiness, runbooks, and on-call practices
  • Influence application teams to adopt shift-left reliability practices
Must-Have Skills & Experience
  • Hands-on experience supporting Java applications in production
  • Strong understanding of JVM fundamentals (heap/memory management, garbage collection, OOM issues, thread analysis)
  • Proven experience with SRE practices, including:
    • Incident response and on-call support
    • Root cause analysis and postmortems
    • SLIs, SLOs, and reliability-driven operations
  • Strong experience troubleshooting using application logs, metrics, and monitoring tools
  • Experience operating Java applications on Kubernetes (EKS) from an application/runtime perspective
  • Experience with deployment strategies (rolling, blue/green, canary)
  • Ability to write automation and scripts (Python or any) to reduce operational toil
  • Solid understanding of application architecture and service dependencies (databases, messaging systems, external APIs)
  • Strong collaboration and communication skills; ability to work closely with development teams
  • Demonstrates accountability and sound judgment when responding to high-pressure incidents
Good-to-Have Skills & Experience
  • Exposure to platform or infrastructure concepts supporting application workloads
  • Experience with AWS services such as EKS, RDS/Aurora, S3, EFS, and CloudWatch
  • CI/CD pipeline experience (GitHub Actions, Jenkins)
  • Familiarity with GitOps practices
  • Experience with cloud migrations or modernization efforts
#J-18808-Ljbffr