1

Linux Site Reliability Engineer Jobs in New York

Staff Site Reliability Engineer

White Plains, NY · On-site

$59 - $78.50/hr

The Site Reliability Engineering team drives reliability strategy, elevates engineering standards ... Deep Linux expertise - from kernel internals and system performance tuning to hardening and ...

Strong Linux and networking fundamentals * Experience operating systems in production environments ... Experience with multi-site or edge deployments * Experience with event-driven systems (Kafka or ...

What a Senior Site Reliability Engineer does at Clover As a Senior Site Reliability Engineer, you ... What you will need to have * 5+ years of production experience in AWS, Kubernetes, and Linux/Unix ...

Site Reliability Engineer

New York, NY · On-site

$62.25 - $82.75/hr

The Role We are seeking highly experienced Site Reliability Engineers (SRE) to shape the reliability, scalability and performance of our platform and customer facing applications. You will work ...

Director, Splunk Platform Engineering & SRE

New York, NY · On-site

$62.25 - $82.75/hr

... & SRE At BNY, our culture allows us to run our company better and enables employees' growth and ... Linux/Unix OS internals (CPU, memory, I/O, process behavior) * Network behavior, packet flow, and ...

Showing results 41-60

Linux Site Reliability Engineer information

What are some common challenges faced by Linux Site Reliability Engineers when scaling infrastructure, and how can they be addressed?

Linux Site Reliability Engineers often encounter challenges related to maintaining system stability and performance as infrastructure scales. Issues such as configuration drift, automation bottlenecks, and monitoring gaps can arise when managing numerous servers or services. Addressing these challenges typically involves implementing robust configuration management tools, investing in automated deployment pipelines, and enhancing observability through comprehensive monitoring and alerting solutions. Collaboration with development and operations teams is essential to ensure that scalability solutions align with business needs and technical requirements.

What are the key skills and qualifications needed to thrive as a Linux Site Reliability Engineer?

To thrive as a Linux Site Reliability Engineer, you need deep expertise in Linux system administration, scripting (such as Bash or Python), and a solid understanding of networking concepts, usually backed by a computer science degree or equivalent experience. Familiarity with configuration management tools (like Ansible, Puppet, or Chef), containerization (Docker, Kubernetes), and cloud platforms (AWS, GCP, or Azure) is typically required, along with relevant certifications like RHCE or AWS Certified SysOps Administrator. Strong problem-solving skills, effective communication, and the ability to work under pressure are crucial soft skills for this role. These competencies ensure the reliability, scalability, and security of complex infrastructure, minimizing downtime and supporting seamless operations.

What is the difference between Linux Site Reliability Engineer vs Linux DevOps Engineer?

AspectLinux Site Reliability EngineerLinux DevOps Engineer
CredentialsLinux certifications, SRE-specific trainingLinux certifications, DevOps tools certifications
Work EnvironmentFocus on system reliability, monitoring, incident responseFocus on automation, CI/CD pipelines, deployment
Employer & IndustryTech companies, cloud providers, large enterprisesStartups, tech firms, software development teams
Search & Comparison IntentUnderstanding reliability roles, incident managementAutomation, deployment, continuous integration

While both roles involve Linux expertise, a Linux Site Reliability Engineer primarily focuses on maintaining system reliability, monitoring, and incident response. In contrast, a Linux DevOps Engineer emphasizes automation, continuous integration, and deployment processes. Both roles require Linux skills and often overlap, but their core responsibilities differ based on organizational needs.

What is a Linux Site Reliability Engineer?

A Linux Site Reliability Engineer (SRE) is an IT professional responsible for ensuring the reliability, scalability, and performance of systems running on the Linux operating system. They bridge the gap between software development and operations by automating processes, monitoring infrastructure, and managing incidents. Linux SREs focus on system availability, building tools for deployment and monitoring, and improving system robustness through best practices and automation. Their work helps organizations deliver reliable online services and quickly recover from outages or system failures.
What are popular job titles related to Linux Site Reliability Engineer jobs in New York? For Linux Site Reliability Engineer jobs in New York, the most frequently searched job titles are:
What job categories do people searching Linux Site Reliability Engineer jobs in New York look for? The top searched job categories for Linux Site Reliability Engineer jobs in New York are:
What cities in New York are hiring for Linux Site Reliability Engineer jobs? Cities in New York with the most Linux Site Reliability Engineer job openings:

Senior Site Reliability Engineer

Castleton Commodities International, LLC

Stamford, CT • On-site

$60.75 - $80.75/hr

Full-time

Medical, Dental, Life, Retirement, PTO

This job post has expired today. Applications are no longer accepted.


Job description

The Senior Site Reliability Engineer is responsible for improving the reliability, availability, scalability, and operational excellence of our critical infrastructure platforms and services. This role partners closely with Engineering, Security, and Infrastructure teams to design resilient cloud-native architectures, implement Infrastructure as Code (IaC) and CI/CD standards, and drive measurable reliability outcomes. The Senior Site Reliability Engineer will also lead efforts to define and validate recovery objectives (RTO/RPO), design and implement Business Continuity / Disaster Recovery (BCP/DR) plans, and coordinate structured testing to ensure readiness.
Responsibilities:
Reliability Engineering & Operations
  • Own and improve service reliability through SLO/SLI definition, error budgets, and operational best practices.
  • Design, implement, and maintain observability (monitoring, logging, tracing, alerting) to reduce MTTR and improve proactive detection.
  • Lead incident response practices including on-call improvements, runbooks, post-incident reviews (RCA), and preventative actions.
  • Partner with application teams to improve performance, capacity planning, and resiliency under failure scenarios.

Infrastructure & Cloud Architecture
  • Design and operate highly available, fault-tolerant Cloud architectures (multi-AZ and, where required, multi-region).
  • Implement resilient patterns across compute, storage, networking, and managed services (e.g., autoscaling, load balancing, backups, replication).
  • Drive cloud governance best practices (tagging, account/landing zone patterns, least privilege, guardrails) in partnership with security and platform teams.

Infrastructure as Code (IaC) & DevOps Enablement
  • Build and maintain IaC modules and standards (e.g., Terraform, CloudFormation, CDK) for repeatable, auditable infrastructure delivery.
  • Develop, standardize, and optimize CI/CD pipelines to enable safe, automated deployments (e.g., GitHub Actions, GitLab CI, Jenkins, AWS CodePipeline).
  • Promote DevOps practices: version-controlled infrastructure, automated testing, immutable deployments, and progressive delivery patterns.
  • Establish environment consistency across dev/test/stage/prod and ensure infrastructure drift detection and remediation.

BCP/DR, RTO/RPO Definition & Testing
  • Collaborate with stakeholders to evaluate and define service-level RTO and RPO targets based on business and technical requirements.
  • Design and implement BCP/DR architectures and procedures (backups, restore workflows, replication, failover/failback, data integrity validation).
  • Coordinate and execute structured DR tests (tabletop, simulation, partial failover, full failover) and document outcomes.
  • Maintain DR runbooks, dependency maps, and recovery checklists; drive remediation of gaps identified during testing.
  • Produce metrics and reporting on DR readiness, test results, and continuous improvement actions.

Qualifications:
  • 7+ years of experience in SRE, DevOps, Platform Engineering, or Systems Engineering roles supporting production environments.
  • Strong proficiency with observability platforms (e.g., Datadog, Prometheus/Grafana, ELK/OpenSearch, Nagios, Nimsoft, etc).
  • Strong hands-on AWS experience building and operating production systems.
  • Proven expertise with Infrastructure as Code (Terraform and/or CloudFormation/CDK).
  • Strong CI/CD and automation background (pipeline design, deployment strategies, testing automation).
  • Experience defining and validating RTO/RPO, and implementing BCP/DR plans with structured testing.
  • Experience with Kubernetes and auto-scaling container platforms (EKS, ECS, or Kubernetes on-prem).
  • Strong Linux fundamentals, networking concepts (DNS, TCP/IP, load balancing), and troubleshooting skills.
  • Proficiency in at least one scripting/programming language (Python, Go, Bash, or similar).
  • Ability to write clear operational documentation, runbooks, and post-incident reports.
  • Ability to work effectively in a fast-paced, dynamic and high-intensity environment including open-floor plan if applicable to the position, with timely responsiveness and the ability to work beyond normal business hours when required.

Preferred Qualifications:
  • Familiarity with Azure and/or Oracle Cloud (OCI).
  • Familiarity with Service Mesh, API Gateways, and distributed tracing tooling.
  • Familiarity with OpenTelemetry, client instrumentations and collector configurations.
  • Security and compliance familiarity in cloud environments (IAM design, secrets management, audit logging).
  • Experience implementing progressive delivery (blue/green, canary), feature flags, and automated rollback.
  • Relevant certifications (AWS Solutions Architect/DevOps Engineer, Kubernetes CKA/CKAD).
  • Experience with ArgoCD & Karpenter.

Employee Programs & Benefits:
CCI offers competitive benefits and programs to support our employees, their families and local communities. These include:
  • Competitive comprehensive medical, dental, retirement and life insurance benefits
  • Employee assistance & wellness programs
  • Parental and family leave policies
  • CCI in the Community: Each office has a Charity Committee and as a part of this program employees are allocated 2 days annually to volunteer at the selected charities.
  • Charitable contribution match program
  • Tuition assistance & reimbursement
  • Quarterly Innovation & Collaboration Awards
  • Employee discount program, including access to fitness facilities
  • Competitive paid time off
  • Continued learning opportunities

Visit https://www.cci.com/careers/life-at-cci/# to learn more!
#LI-CD1