1

Network Reliability Engineer Jobs in Oakbrook Terrace, IL

Sr. Site Reliability Engineer (SRE)

Chicago, IL · On-site

$58.75 - $78/hr

The Sr. Site Reliability Engineer will be responsible for building and operating production-grade ... Configure CNI plugins and network segmentation for research workloads. • Develop and maintain ...

Staff Site Reliability Engineer (SRE) - Platform Engineering Note: This position follows a hybrid ... Deep-seated expertise in GCP (Networking, IAM, GKE) and the ability to scale Kafka clusters for ...

Staff Site Reliability Engineer (SRE) - Platform Engineering Note: This position follows a hybrid ... Deep-seated expertise in Google Cloud Platform (Networking, IAM, GKE) and the ability to scale ...

Staff Site Reliability Engineer (SRE) - Platform Engineering Note: This position follows a hybrid ... Deep-seated expertise in GCP (Networking, IAM, GKE) and the ability to scale Kafka clusters for ...

Sr. Site Reliability Engineer (SRE)

Chicago, IL · On-site

$58.75 - $78/hr

Configure CNI plugins and network segmentation for research workloads. * Custom Operators ... Experience: 5+ years in SRE, DevOps, or infrastructure engineering roles with proven experience ...

Site Reliability Engineer

Chicago, IL · On-site

$155K - $222K/yr

Meet the Team The SRE Fleet team is responsible for maintaining the stability, scalability, and ... Add to that our worldwide network of doers and experts, and you'll see that the opportunities to ...

Configure CNI plugins and network segmentation for research workloads. * Custom Operators ... Experience: 5+ years in SRE, DevOps, or infrastructure engineering roles with proven experience ...

Fitch Group is currently seeking a Senior Service Reliability Engineer to embed with Fitch ... networking, storage, DNS) and APM/telemetry tooling. What Would Make You Stand Out: * Practical ...

As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and Investment Banking ... Implements infrastructure, configuration, and network as code for the applications and platforms in ...

New

Site Reliability Engineer III

Chicago, IL · On-site

$58.75 - $78/hr

As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and Investment Banking ... Implements infrastructure, configuration, and network as code for the applications and platforms in ...

As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and Investment Banking ... Implements infrastructure, configuration, and network as code for the applications and platforms in ...

New

next page

Showing results 1-20

Network Reliability Engineer information

See Oakbrook Terrace, IL salary details

$61.3K

$118.6K

$141.8K

How much do network reliability engineer jobs pay per year?

As of Aug 7, 2026, the average yearly pay for network reliability engineer in Oakbrook Terrace, IL is $118,639.00, according to ZipRecruiter salary data. Most workers in this role earn between $103,100.00 and $129,700.00 per year, depending on experience, location, and employer.

What is a network reliability engineer?

A Network Reliability Engineer (NRE) is an IT professional responsible for ensuring the reliability, performance, and scalability of network systems. They combine skills in networking, software engineering, and automation to proactively detect and resolve potential network issues before they affect users. NREs often design and implement monitoring tools, automate network management tasks, and work to improve the overall stability of network infrastructure. Their goal is to minimize downtime and ensure seamless connectivity across an organization’s network.

What is the difference between Network Reliability Engineer vs Network Operations Center (NOC) Technician?

AspectNetwork Reliability EngineerNetwork Operations Center (NOC) Technician
CertificationsCCNA, CCNP, Network+CCNA, Network+
Work EnvironmentDesign, analyze, and improve network infrastructureMonitor, troubleshoot, and maintain networks in real-time
Employer & Industry UsageTelecom, large enterprises, cloud providersISPs, data centers, enterprise networks
Common Search & ComparisonFocus on network reliability and designFocus on network monitoring and incident response

The main difference is that Network Reliability Engineers focus on designing and improving network systems to ensure long-term reliability, while NOC Technicians monitor and troubleshoot networks in real-time to resolve issues quickly. Both roles require relevant certifications and are essential in maintaining network performance, but they serve different functions within network management.

What are the key skills and qualifications needed to thrive as a network reliability engineer, and why are they important?

To thrive as a Network Reliability Engineer, you need a strong background in computer networking, network protocols, troubleshooting, and often a degree in computer science or a related field. Familiarity with tools like Wireshark, Nagios, Cisco IOS, and certifications such as CCNA or CCNP are commonly required. Analytical thinking, proactive problem-solving, and effective communication are standout soft skills in this role. These skills are crucial to maintaining reliable network operations, minimizing downtime, and ensuring seamless communication across organizational systems.

What are some common challenges faced by network reliability engineers, and how are they typically addressed?

Network Reliability Engineers often encounter challenges such as diagnosing intermittent connectivity issues, managing network upgrades with minimal downtime, and maintaining high availability during peak traffic. These are typically addressed by leveraging robust monitoring tools, implementing automation for routine tasks, and collaborating closely with software engineers, network administrators, and incident response teams. Staying current with evolving network technologies and best practices is essential for effectively identifying and resolving problems before they impact users.
What cities near Oakbrook Terrace, IL are hiring for Network Reliability Engineer jobs? Cities near Oakbrook Terrace, IL with the most Network Reliability Engineer job openings:
Infographic showing various Network Reliability Engineer job openings in Oakbrook Terrace, IL as of July 2026, with employment types broken down into 1% As Needed, 77% Full Time, 14% Part Time, 7% Contract, and 1% Nights. Highlights an 92% Physical, 2% Hybrid, and 6% Remote job distribution, with an average salary of $118,639 per year, or $57 per hour.

Sr. Site Reliability Engineer (SRE)

Moonlite AI

Chicago, IL • On-site

$58.75 - $78/hr

Full-time

Re-posted 24 days ago


Job description

Job Summary:
Moonlite AI delivers high-performance AI infrastructure for organizations running intensive computational research and large-scale model training. The Sr. Site Reliability Engineer will be responsible for building and operating production-grade AI infrastructure, focusing on Kubernetes expertise to ensure reliability and performance across multiple regions.
Responsibilities:
• Design, build, and operate production Kubernetes clusters on bare-metal infrastructure – including cluster bootstrapping, control plane architecture, etcd management, and scaling strategies for high-performance compute workloads.
• Implement and operate custom Kubernetes networking solutions with SR-IOV for high-performance GPU interconnects, multi-tenancy isolation and advanced networking policies. Configure CNI plugins and network segmentation for research workloads.
• Develop and maintain custom Kubernetes operators and controllers for bare-metal provisioning, infrastructure lifecycle management, and resource orchestration across compute, storage, and networking domains.
• Deploy and optimize NVIDIA GPU operators, device plugins, and other custom scheduling logic for GPU workload placement and utilization optimization.
• Build deep integrations between Kubernetes and underlying infrastructure including CSI drivers for storage, custom admission controllers for policy enforcement, and scheduling extensions for specialized hardware placement.
• Design and implement automation using Terraform, Ansible, Helm, and custom operators to orchestrate infrastructure workflows and enable deployments across multiple regions.
• Manage production bare-metal infrastructure across multiple regions. Build systems ensuring high availability, fault tolerance, and graceful degradation – establishing SLIs, SLOs, and monitoring to meet enterprise reliability commitments.
• Build comprehensive monitoring, logging, and alerting using Prometheus, Grafana, and ELK stack. Lead incident response, conduct postmortems, and implement preventative measures to improve reliability and reduce MTTR.
• Identify and resolve performance bottlenecks across infrastructure domains. Monitor utilization trends, forecast capacity needs, and optimize resource allocation for various workloads.
Qualifications:
Required:
• 5+ years in SRE, DevOps, or infrastructure engineering roles with proven experience operating production infrastructure at scale.
• Deep hands-on experience building and operating production Kubernetes clusters on bare-metal infrastructure – not just deploying workloads in managed clusters. Must understand cluster bootstrapping, control plane architecture, etcd operations, and scaling strategies.
• Strong understanding of Kubernetes internals including custom resource definitions (CRDs), operators, controllers, admission webhooks, and scheduling. Experience integrating storage (CSI drivers), networking (CNI, SR-IOV), and specialized hardware (GPU device plugins) with Kubernetes.
• Strong fundamentals in Linux systems administration, performance tuning, troubleshooting, and automation in production environments.
• Proficiency with infrastructure-as-code tools (Terraform, Ansible, Helm) and building automation to reduce operational overhead.
• Solid understanding of networking concepts including IPAM, DNS, DHCP, VLAN/VXLAN, routing, load balancing, and experience troubleshooting network issues in production.
• Experience building and maintaining comprehensive monitoring solutions using tools like Prometheus, Grafana, and centralized logging systems.
• Understanding of SRE principles including SLIs/SLOs/SLAs, error budgets, incident management, and blameless postmortems.
• Strong scripting skills in Go, Python, or Bash for automation, tooling development, and operational efficiency.
• Demonstrated ability to troubleshoot complex issues under pressure, manage incidents effectively, and communicate clearly during outages.
• Excellent communication skills and ability to work across teams including systems engineers, network engineers, and software developers.
Preferred:
• Experience building custom Kubernetes operators or controllers for infrastructure orchestration
• Deep familiarity with Kubernetes networking (Calico, Cilium, Multus), service mesh technologies, and network policy management
• Experience with GPU workload orchestration including NVIDIA GPU Operator, MIG, time-slicing, and device plugins
• Background with advanced Kubernetes features including custom schedulers, admission controllers, and API server extensions
• Experience with Kubernetes cluster federation or multi-cluster management
• Knowledge of high-performance networking technologies (InfiniBand, RDMA, RoCE) and their integration with Kubernetes
• Experience with enterprise storage systems (VAST, Lightbits, Ceph, or similar)
• Familiarity with configuration management at scale and GitOps practices
• Understanding of security best practices for Kubernetes and bare-metal infrastructure
• Experience operating infrastructure in regulated industries or co-located data center environments
• Background supporting research institutions, technical computing environments, or enterprise AI infrastructure
Company:
Moonlite AI is a technology company. Founded in 2024, the company is headquartered in Chicago, USA, with a team of 2-10 employees. The company is currently Early Stage.