1

Network Observability Jobs in Chicago, IL (NOW HIRING)

Data Engineering Lead (Hybrid)

Chicago, IL · On-site +1

$105K - $139K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

About Rewards Network For 41 years, Rewards Network has been helping restaurants grow revenue ... Own data pipeline monitoring and observability , ensuring production pipelines have appropriate ...

Data Engineering Lead (Hybrid)

Chicago, IL · Hybrid

$105K - $139K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

About Rewards Network For 41 years, Rewards Network has been helping restaurants grow revenue ... Own data pipeline monitoring and observability , ensuring production pipelines have appropriate ...

Data Engineering Lead (Hybrid)

Chicago, IL · Hybrid

$105K - $139K/yr

About Rewards Network For 41 years, Rewards Network has been helping restaurants grow revenue ... Own data pipeline monitoring and observability , ensuring production pipelines have appropriate ...

Senior DevOps Engineer

Chicago, IL

$133K - $172K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Cluster lifecycle management via Cluster API, fleet-wide upgrades, bare metal provisioning, CNI networking, storage, autoscaling, and disaster recovery planning. * Observability: Operate and scale ...

Senior DevOps Engineer

Chicago, IL · On-site

$133K - $172K/yr

Cluster lifecycle management via Cluster API, fleet-wide upgrades, bare metal provisioning, CNI networking, storage, autoscaling, and disaster recovery planning. * Observability: Operate and scale ...

Senior DevOps Engineer

Chicago, IL · On-site

$133K - $172K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Cluster lifecycle management via Cluster API, fleet-wide upgrades, bare metal provisioning, CNI networking, storage, autoscaling, and disaster recovery planning. * Observability: Operate and scale ...

Senior DevOps Engineer

Chicago, IL · On-site

$133K - $172K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Cluster lifecycle management via Cluster API, fleet-wide upgrades, bare metal provisioning, CNI networking, storage, autoscaling, and disaster recovery planning. * Observability: Operate and scale ...

Platform Engineer

Lincolnshire, IL · On-site

$140 - $180/hr

Experience with observability platforms (e.g., Dynatrace, Datadog, Prometheus/Grafana stack) and a solid understanding of the three pillars of observability. * Strong understanding of networking ...

New

Senior Product Manager, Network Platform (Remote)

Chicago, IL · On-site +1

$169K - $237K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

As a core component of Cisco's Networking, Security, and Observability strategy, our work ensures that customers can operate their infrastructure with absolute confidence. We are a fast-moving ...

New

IT Site Reliability & Performance Engineer

Wood Dale, IL · On-site

$110K - $130K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Strong knowledge of monitoring and observability tools and platforms. * Experience with IT infrastructure and application technologies, including Linux, Windows, Azure, networking, databases, and web ...

IT Site Reliability & Performance Engineer

Wood Dale, IL · On-site

$110K - $130K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Strong knowledge of monitoring and observability tools and platforms. * Experience with IT infrastructure and application technologies, including Linux, Windows, Azure, networking, databases, and web ...

Sr. Site Reliability Engineer (SRE)

Chicago, IL · On-site

$58.75 - $78/hr

Solid understanding of networking concepts including IPAM, DNS, DHCP, VLAN/VXLAN, routing, load balancing, and experience troubleshooting network issues in production. * Observability & Monitoring:

Python, Go, or Java;plusBash and solid Linux fundamentals (networking, filesystems, JVM tuning basics). * Observability and reliability engineering for Kafka: Prometheus/Grafana, logging, alerting ...

New

Showing results 41-60

Network Observability information

What is network observability?

Network observability refers to the ability to gain deep visibility into all aspects of a computer network’s operations, performance, and health. It involves collecting and analyzing data from various sources such as logs, metrics, and traces to detect issues, optimize performance, and ensure security. Unlike traditional monitoring, observability provides insights into the underlying causes of network problems, enabling faster troubleshooting and proactive management. Network observability tools help organizations maintain reliable, secure, and efficient network infrastructure, which is especially important in complex or large-scale environments.

What are the key skills and qualifications needed to thrive in network observability, and why are they important?

To excel in Network Observability, you need a strong understanding of networking fundamentals, troubleshooting, and data analysis, often supported by a degree in computer science or a related field. Familiarity with monitoring tools like Wireshark, Datadog, Grafana, and knowledge of protocols such as SNMP and NetFlow, as well as relevant certifications (e.g., Cisco CCNA/CCNP), are typically required. Strong problem-solving, attention to detail, and communication skills help professionals quickly identify issues and work with cross-functional teams. These competencies are essential for ensuring network reliability, performance, and rapid incident response in complex IT environments.

What are some common challenges faced by professionals in network observability roles, and how can they be addressed?

Professionals in Network Observability often face challenges like managing large volumes of network data, integrating diverse monitoring tools, and quickly identifying the root cause of network issues. Addressing these challenges typically involves automating data collection, leveraging advanced analytics, and fostering close collaboration with network engineers and security teams. Staying updated with the latest observability platforms and best practices can also help streamline workflows and improve network performance monitoring.

What is the difference between Network Observability vs Network Monitoring?

AspectNetwork ObservabilityNetwork Monitoring
FocusProvides comprehensive insights into network performance, health, and security through data collection, analysis, and visualization.Tracks network uptime, availability, and basic performance metrics to detect outages or issues.
Tools & SkillsUses advanced analytics, telemetry, and machine learning; requires knowledge of data analysis and network architecture.Utilizes monitoring tools like SNMP, ping, and simple dashboards; requires basic network troubleshooting skills.
PurposeEnables proactive troubleshooting, capacity planning, and security analysis by understanding complex network behaviors.Provides real-time alerts and status updates to quickly identify and resolve network issues.

While both roles focus on network health, Network Observability offers a deeper, data-driven understanding of network behavior, supporting proactive management. Network Monitoring is more about real-time detection of outages and basic performance tracking. Organizations often use both to ensure robust network performance and security.

Is network observability a good career?

Network observability is a growing field that involves monitoring and analyzing network performance using tools like telemetry, logs, and metrics. It requires skills in networking, scripting, and understanding of network infrastructure, making it a valuable and in-demand career path with opportunities for advancement. The role often involves working with cloud environments and automation tools, and certifications such as Cisco or Cisco DevNet can enhance job prospects.

What are the most commonly searched types of Network Observability jobs in Chicago, IL?

The most popular types of Network Observability jobs in Chicago, IL are:

What are popular job titles related to Network Observability jobs in Chicago, IL?

For Network Observability jobs in Chicago, IL, the most frequently searched job titles are:

What job categories do people searching Network Observability jobs in Chicago, IL look for?

The top searched job categories for Network Observability jobs in Chicago, IL are:

What cities near Chicago, IL are hiring for Network Observability jobs?

Cities near Chicago, IL with the most Network Observability job openings:

Platform Engineering & AI Operations Lead

Intelligent Generation

Oak Brook, IL • On-site

$103K - $136K/yr

Full-time

Medical, Retirement, PTO

Re-posted 11 days ago


Job description

Benefits:
  • Warrants
  • Annual bonus
  • 401(k)
  • Competitive salary
  • Health insurance
  • Paid time off

Platform Engineering & AI Operations Lead
Full Time | Hybrid | Chicago Metro Area

Build the platform foundation for POWR:Suite and IG’s AI-assisted operating model
Intelligent Generation’s mission is to empower businesses to engage the clean energy grid. 
Intelligent Generation builds and operates POWR:Suite, a software platform that helps battery energy storage assets make highly profitable economic decisions.
POWR:Suite connects distributed energy assets to wholesale power markets while also optimizing behind-the-meter value: reducing utility bills, managing demand charges, improving asset performance, supporting resilience, and helping customers capture the full economic value of their energy assets.
Our work sits at the intersection of energy markets, grid operations, customer savings, software automation, telemetry, and AI-assisted decision-making.
We are looking for a platform engineering leader who can own the infrastructure foundation behind POWR:Suite and help IG scale an AI-assisted engineering and operating model.
This is a leadership role. You will be hands-on early, but the expectation is that you will grow into leading people, establishing platform standards, and orchestrating AI agents that improve engineering delivery, reliability, security, and operations.
Why this role matters
Battery assets are only valuable when they are operated intelligently. Every decision matters: when to charge, when to discharge, when to participate in the market, when to preserve state of charge, when to reduce customer demand charges, when to support resilience, and how to prove the economic value created.
The platform behind POWR:Suite must be reliable, observable, secure, automated, and ready to support both human operators and AI-assisted workflows.
You will help define that foundation.

What you will lead and own
Cloud platform and infrastructure
Lead the architecture and operation of IG’s cloud platform on GCP, including Cloud Run, GKE, Cloud SQL, Pub/Sub, IAM, Terraform, networking, observability, and deployment standards.
CI/CD and engineering enablement
Build and lead the practices that make engineering delivery faster, safer, and more repeatable. Own CI/CD, release automation, environment strategy, deployment quality, and platform guardrails.
Edge-to-cloud reliability
Strengthen the communication layer between field assets and cloud systems. Partner with software, controls, and operations teams to improve telemetry, command paths, failure detection, and recovery patterns.
Security and operational controls
Own practical security and IT operations standards for a company operating live energy infrastructure, including IAM lifecycle, endpoint standards, secrets management, access controls, and production system hygiene.
Agent operations foundation
Build, maintain, evaluate, and govern agents that support platform engineering work, including deployment assistants, infrastructure review agents, security review agents, incident summarizers, documentation agents, and operational troubleshooting agents.
People and agent orchestration
Over time, build and lead a platform engineering function. Establish how work is divided between engineers and agents, how agent outputs are reviewed, and how platform knowledge compounds over time.

What success looks like
First 90 days

  • Understand current platform architecture, cloud services, deployments, edge/cloud interfaces, and operational pain points
  • Establish a platform ownership map and priority risk list
  • Improve documentation around environments, deployment flow, access, and operational dependencies
  • Identify high-value opportunities for agent-assisted platform work
First 6 months
  • Improve CI/CD maturity and release reliability
  • Strengthen observability and alerting for critical platform services
  • Build or deploy early internal agents that assist with platform review, troubleshooting, documentation, or deployment support
  • Establish platform standards that engineering teams can follow

First 12 months

  • Lead the platform function for POWR:Suite
  • Improve engineering independence, reliability, and operational control
  • Mature agent-assisted engineering workflows
  • Build the foundation for a team that can scale with IG’s growth
What we are looking for
Required

  • 8+ years in platform engineering, infrastructure engineering, DevOps, SRE, cloud architecture, or senior software engineering
  • Experience leading technical work across teams or mentoring engineers
  • Strong production GCP experience
  • Terraform or infrastructure-as-code experience
  • CI/CD ownership in production environments
  • Strong Python or equivalent engineering/scripting ability
  • Strong understanding of distributed systems, reliability, observability, and operational failure modes
  • Security-first mindset around IAM, access control, secrets, production systems, and operational risk
  • Hands-on experience using AI tools as part of engineering work
  • Ability to build, maintain, evaluate, govern, or orchestrate agents that support engineering workflows
  • Strong systems thinking and ownership mindset
Strongly preferred
  • Energy, industrial, IoT, SCADA, or OT-adjacent experience
  • Familiarity with DNP3, Modbus, ICCP, or similar protocols
  • Experience with agentic AI systems, tool use, evals, or AI-assisted operations
  • GCP Professional certification
  • Experience building or leading a platform engineering team
Why this is exciting
You will help define the platform foundation for a company operating at the intersection of energy storage, virtual power plants, cloud infrastructure, market automation, customer savings, and AI-assisted engineering.
This is a chance to build systems, standards, agents, and eventually a team that will shape how IG scales POWR:Suite and optimizes the full economic value of distributed battery assets.
This position is hybrid and includes both remote and in-person work.
Intelligent Generation participates in the E-Verify process for all new hires.

Flexible work from home options available.