1

Linux Site Reliability Engineer Jobs in New York

DNS Engineer - SRE

Bethpage, NY

$58.25 - $77.50/hr

Perform Linux kernel tuning for high-performance network throughput and conduct deep-dive log ... SRE principles (automation, reliability, and monitoring) in production environments DNS ...

DNS Engineer - SRE

Bethpage, NY · On-site

$58.25 - $77.50/hr

Perform Linux kernel tuning for high-performance network throughput and conduct deep-dive log ... SRE principles (automation, reliability, and monitoring) in production environments • DNS ...

DNS Engineer - SRE

Bethpage, NY · On-site

$58.25 - $77.50/hr

Job Summary The Role DNS Engineer - SRE is a high-impact role responsible for the architecture ... Perform Linux kernel tuning for high-performance network throughput and conduct deep-dive log ...

Site Reliability Engineer

Bethpage, NY · Hybrid

$58.25 - $77.50/hr

Job Summary As a Site Reliability Engineer II, you will be a primary driver in the long-term ... Audit, harden, and standardize Unix (Solaris/AIX) and Linux (RHEL/Ubuntu) environments across both ...

Site Reliability Engineer

New York, NY · On-site +1

$120K - $160K/yr

As a Site Reliability Engineer (SRE), you will work at the intersection of production operations ... Experience with Linux * Familiarity with Relational Databases & SQL * Sharp analytical and problem ...

Application SRE (DevOps)

Elmwood Park, NJ · On-site

$100K - $120K/yr

Solid understanding of Windows, Linux/Unix systems and networking fundamentals * 7 years of experience as an SRE * Hands-on experience with cloud platforms such as AWS, Azure, or GCP * Experience ...

Site Reliability Engineer

New York, NY · On-site

$62.25 - $82.75/hr

As a Site Reliability Engineer you will be driving the entire Justworks infrastructure. In your ... We are happy to relocate the right candidate Technologies We Use AWS, Linux, Puppet, Memcached ...

Email Reliability Engineer

New York, NY · On-site

$62.25 - $82.75/hr

Senior Systems Engineer - Email Reliability (Hybrid SRE) New York, NY (Hybrid, 3 days in office ... You will own the reliability and architecture of a complex hybrid stack (Linux/Postfix + Microsoft ...

Site Reliability Engineer

New York, NY · On-site

$62.25 - $82.75/hr

As a Site Reliability Engineer you will be driving the entire Justworks infrastructure. In your ... We are happy to relocate the right candidate Technologies We Use AWS, Linux, Puppet, Memcached ...

Site Reliability Engineer, Pragma

New York, NY · Hybrid

$62.25 - $82.75/hr

The Role We are seeking a Site Reliability Engineer to join a team responsible for the production ... Understanding of JVM-based distributed applications and experience across Linux systems. * Solid ...

Site Reliability Engineer, Pragma

New York, NY · On-site

$62.25 - $82.75/hr

The Role We are seeking a Site Reliability Engineer to join a team responsible for the production ... Understanding of JVM-based distributed applications and experience across Linux systems. * Solid ...

As one of our Lead Site Reliability Engineers, you will combine hands-on technical expertise with ... Advanced Linux systems expertise, with the ability to diagnose complex system-level issues and ...

Showing results 21-40

Linux Site Reliability Engineer information

What are some common challenges faced by Linux Site Reliability Engineers when scaling infrastructure, and how can they be addressed?

Linux Site Reliability Engineers often encounter challenges related to maintaining system stability and performance as infrastructure scales. Issues such as configuration drift, automation bottlenecks, and monitoring gaps can arise when managing numerous servers or services. Addressing these challenges typically involves implementing robust configuration management tools, investing in automated deployment pipelines, and enhancing observability through comprehensive monitoring and alerting solutions. Collaboration with development and operations teams is essential to ensure that scalability solutions align with business needs and technical requirements.

What are the key skills and qualifications needed to thrive as a Linux Site Reliability Engineer?

To thrive as a Linux Site Reliability Engineer, you need deep expertise in Linux system administration, scripting (such as Bash or Python), and a solid understanding of networking concepts, usually backed by a computer science degree or equivalent experience. Familiarity with configuration management tools (like Ansible, Puppet, or Chef), containerization (Docker, Kubernetes), and cloud platforms (AWS, GCP, or Azure) is typically required, along with relevant certifications like RHCE or AWS Certified SysOps Administrator. Strong problem-solving skills, effective communication, and the ability to work under pressure are crucial soft skills for this role. These competencies ensure the reliability, scalability, and security of complex infrastructure, minimizing downtime and supporting seamless operations.

What is the difference between Linux Site Reliability Engineer vs Linux DevOps Engineer?

AspectLinux Site Reliability EngineerLinux DevOps Engineer
CredentialsLinux certifications, SRE-specific trainingLinux certifications, DevOps tools certifications
Work EnvironmentFocus on system reliability, monitoring, incident responseFocus on automation, CI/CD pipelines, deployment
Employer & IndustryTech companies, cloud providers, large enterprisesStartups, tech firms, software development teams
Search & Comparison IntentUnderstanding reliability roles, incident managementAutomation, deployment, continuous integration

While both roles involve Linux expertise, a Linux Site Reliability Engineer primarily focuses on maintaining system reliability, monitoring, and incident response. In contrast, a Linux DevOps Engineer emphasizes automation, continuous integration, and deployment processes. Both roles require Linux skills and often overlap, but their core responsibilities differ based on organizational needs.

What is a Linux Site Reliability Engineer?

A Linux Site Reliability Engineer (SRE) is an IT professional responsible for ensuring the reliability, scalability, and performance of systems running on the Linux operating system. They bridge the gap between software development and operations by automating processes, monitoring infrastructure, and managing incidents. Linux SREs focus on system availability, building tools for deployment and monitoring, and improving system robustness through best practices and automation. Their work helps organizations deliver reliable online services and quickly recover from outages or system failures.
What are popular job titles related to Linux Site Reliability Engineer jobs in New York? For Linux Site Reliability Engineer jobs in New York, the most frequently searched job titles are:
What job categories do people searching Linux Site Reliability Engineer jobs in New York look for? The top searched job categories for Linux Site Reliability Engineer jobs in New York are:
What cities in New York are hiring for Linux Site Reliability Engineer jobs? Cities in New York with the most Linux Site Reliability Engineer job openings:

$58.25 - $77.50/hr

Full-time

Re-posted 26 days ago


Optimum rating

7.5

Company rating: 7.5 out of 10

Based on 54 frontline employees who took The Breakroom Quiz

48th of 97 rated telecommunications companies


Job description

Are you looking to Optimize your life? Start your exciting path to a rewarding career today!

  
We are Optimum, a leader in the fast-paced world of connectivity, and we're seeking driven and enthusiastic professionals to join our team, empower lives, fuel businesses, and drive innovation. Connectivity is now longer a luxury, but a necessity. A career at Optimum means you'll be enabling progress and enhancing lives by providing reliable, high-speed connectivity solutions that keep the world connected. Our successes, now and in the future, are powered by our amazing product, a commitment to our people and culture, and the connections we make in our communities.


If you are resourceful, collaborative, and passionate about delivering consistent excellence, Optimum is for you! 

Job Summary

 
The Role DNS Engineer - SRE is a high-impact role responsible for the architecture, scalability, and reliability of the mission-critical DNS infrastructure powering our ISP and core network services. This position is designed for an engineer who views infrastructure through the lens of Site Reliability Engineering (SRE) prioritizing automation, observability, and self-healing systems over manual intervention. You will combine deep IP networking and DNS expertise with modern security protocols to ensure our platforms remain resilient against evolving threats and perform at the highest level for millions of users.
The Impact This is a collaborative and influential role. Beyond core engineering, you will serve as a technical authority, leading cross-functional initiatives with Product, Security, and Service Assurance teams. Your goal is to deliver a carrier-grade DNS ecosystem that balances cutting-edge privacy standards (DoH/DoT) with the uncompromising availability required by Tier-1 network operations.

Responsibilities

Core Platform Strategy & Leadership 
   Architectural Ownership: Lead the design and evolution of global DNS architectures, ensuring high availability through Anycast routing, multi-provider redundancy, and automated failover mechanisms.
   Strategic Vendor Relations: Act as the primary technical authority in engagements with DNS and infrastructure vendors, driving roadmaps that align with our long-term reliability and security goals.
   Lifecycle & Capacity Management: Oversee the full lifecycle of DNS platforms-including automated software deployments, hardware refreshes, and proactive capacity planning-to stay ahead of traffic growth.
   Standardization & Policy: Optimize, Define and enforce organization-wide standards for DNS record management, security protocols (DNSSEC), and traffic steering policies to optimize user latency.
   Reliability Engineering: Convert "Strategic Design" into "Operational Reality" by defining Service Level Objectives (SLOs) and Error Budgets for all core name services.
Cross-Domain DNS Operations & SRE
   Protocol Management: Manage the nuances of UDP/TCP port 53, recursion vs. iteration, and complex record types (A, AAAA, CNAME, MX, TXT, SRV).
   Security & Mitigation: Implement and manage DNSSEC to prevent cache poisoning; act as a subject matter expert in mitigating DDoS and DNS amplification attacks.
   Automation (Eliminating Toil): Replace manual updates and "pool" management with automated workflows using Python, Go, Ansible, or Terraform.
   Performance Tuning: Perform Linux kernel tuning for high-performance network throughput and conduct deep-dive log analysis on systems like BIND, Unbound, or PowerDNS.
   Observability: Utilize Prometheus, Grafana, and dnstap to monitor query rates and latency, providing actionable insights into error codes (NXDOMAIN, SERVFAIL).

Qualifications

Minimum Qualifications
   Education: Bachelor's degree in Computer Science, Telecommunications, or a related field (or equivalent practical experience in networking and security)
   Experience: 5+ years in a networking or systems engineering role, with a focus on SRE principles (automation, reliability, and monitoring) in production environments
   DNS Fundamentals: Hands-on experience configuring and maintaining at least two of the following: BIND, Unbound, PowerDNS, AWS Route 53, or Azure DNS
   Networking Protocols: Functional understanding of TCP/IP (IPv4/v6) and DNS-specific protocols including DNSSEC and encrypted transport (DoH/DoT)
   Systems & Automation: Strong Linux/Unix administration skills and proficiency in at least one scripting language (Python, Bash, or Go) for task automation
   Observability: Experience using Grafana and OpenTelemetry (or similar tools) to monitor service health and performance
Preferred Qualifications
   DNS Systems: Hands-on experience managing BIND, Unbound, or PowerDNS in high-traffic environments, alongside cloud-native solutions (AWS Route 53, Azure DNS, Google Cloud DNS)
   Protocol Expertise: Mastery of DNS-specific protocols including DNSSEC, DoT, and DoH, with a firm grasp of underlying transport layers (UDP/TCP) and dual-stack (IPv4/IPv6) networking
   Observability: Experience building dashboards and alerts using Prometheus, ELK, or OpenTelemetry to monitor DNS query latency and error rates
   Automation: Proven ability to manage "DNS as Code" using Terraform or Ansible and writing scripts (Python/Go) to automate routine zone updates
   Scale & Security: Background in Tier-1/Tier-2 service provider environments with a focus on service resilience, Anycast distribution, and DDoS protection
Working Conditions 
   Hybrid remote/on-site, with participation in 24/7 on-call rotations 
   Availability for after-hours maintenance and urgent service restoration activities
   Ability to work in high-pressure, high-reliability production environments

The Ideal Candidate 
The ideal candidate is a proactive engineer who values precision and operational excellence. You don't just manage systems; you architect for reliability, anticipating bottlenecks before they impact the user. We are looking for someone who balances deep technical mastery with an organized approach to delivery, consistently driving improvements in performance, monitoring, and overall service resilience. 

At Optimum, every action and interaction we take part in, is driven by our three Guiding Principles: Do What's Right, Drive One Optimum, and Make It Happen. These aren't just words, they help us build trust, create real community, and embrace new ways of thinking. Our employees are empowered to do the right thing for our customers and co-workers and to recognize and reward these behaviors when we see them. It's all part of the bigger picture of "Be The Difference" where each employee knows they have the power to enact real change, share new ideas, and understand that learning never stop.

If you have the drive to succeed and are ready to embark on a thrilling career, seize this opportunity today, and join our winning team. Together, we'll shape the future of connectivity.

All job descriptions and required skills, qualifications and responsibilities for a particular position are subject to modification by the Company from time to time, in the Company's discretion based on business necessity.

We are an Equal Opportunity Employer committed to recruiting, hiring and promoting qualified people of all backgrounds regardless of gender, race, color, creed, national origin, religion, age, marital status, pregnancy, physical or mental disability, sexual orientation, gender identity, military or veteran status, or any other basis protected by federal, state, or local law.

The Company collects personal information about its applicants for employment that may include personal identifiers, professional or employment related information, photos, education information and/or protected classifications under federal and state law. This information is collected for employment purposes, including identification, work authorization, FCRA-compliant background screening, human resource administration and compliance with federal, state and local law.

Applicants for employment with The Company will never be asked to provide money (even if reimbursable) as part of the job application or hiring process. Please review our Fraud FAQ for further details.

Pay is competitive and based on a number of job-related factors, including skills and experience. The starting pay rate/range at time of hire for this position in New York is $83,538.00 - $137,241.00 / year. For other locations, please inquire with your recruiter. The rates/ranges provided herein are the anticipated pay at the time of hire, and do not reflect future job opportunity.


What Optimum employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom