1

Observability Manager Jobs in Tennessee (NOW HIRING)

Principal Platform Software Engineer

Nashville, TN · On-site

$130K - $174K/yr

... observability for effective debugging, sharing work with manager and/or lead upon completion. * Conducts performance profiling and optimization of coding, building scalable solutions, and ...

Principal Platform Software Engineer

Nashville, TN · On-site

$130K - $174K/yr

... observability for effective debugging, sharing work with manager and/or lead upon completion. * Conducts performance profiling and optimization of coding, building scalable solutions, and ...

You understand observability deeply and have strong opinions about what makes systems debuggable. You've managed ElasticSearch at scale and know the difference between logs that help and logs that ...

Principal Platform Software Engineer

Nashville, TN · On-site

$130K - $174K/yr

... observability for effective debugging, sharing work with manager and/or lead upon completion. * Conducts performance profiling and optimization of coding, building scalable solutions, and ...

Principal Platform Software Engineer

Nashville, TN · On-site

$130K - $174K/yr

... observability data. The APM platform operates at substantial scale, processing millions of spans ... Define and manage telemetry schemas, versioning, data contracts, migrations, and data-quality ...

Principal Platform Software Engineer

Nashville, TN · On-site

$130K - $174K/yr

... observability for effective debugging, sharing work with manager and/or lead upon completion. * Conducts performance profiling and optimization of coding, building scalable solutions, and ...

... observability, resilience, backup, disaster recovery, and platform lifecycle management. * Oversee and manage all phases of the cloud infrastructure product lifecycle, from concept and discovery ...

... observability, resilience, backup, disaster recovery, and platform lifecycle management. * Oversee and manage all phases of the cloud infrastructure product lifecycle, from concept and discovery ...

... observability, resilience, backup, disaster recovery, and platform lifecycle management. * Oversee and manage all phases of the cloud infrastructure product lifecycle, from concept and discovery ...

The Product Management team is looking for aProduct Manager to lead the roadmap and success of ... observability,efficiencyand performance standards. * Be well versed in technical aspects of cloud ...

Systems Engineer - Cloud Ops

Memphis, TN · On-site

$54.25 - $72.25/hr

RBAC, Pod Security Standards, Secrets management, Workload Identity * Understanding of Kubernetes observability: metrics-server, kubectl top, container resource monitoring * Experience debugging ...

Senior Principal SDE

Nashville, TN · On-site

$121K - $167K/yr

Partner with Product Managers and leadership to translate requirements into roadmaps and execution ... You will set expectations for operational excellence , observability, and proactive service ...

Showing results 21-40

Observability Manager information

What is the difference between Observability Manager vs Site Reliability Engineer?

AspectObservability ManagerSite Reliability Engineer
CredentialsTypically requires experience in monitoring, logging, and cloud tools; certifications like AWS, Google Cloud, or Kubernetes are commonRequires strong background in systems engineering, scripting, and cloud platforms; certifications like AWS, GCP, or Linux are often preferred
Work EnvironmentFocuses on overseeing observability tools, data analysis, and team coordination in tech environmentsHands-on role involving system automation, incident response, and infrastructure reliability
Industry UsageUsed across tech companies to improve system visibility and performanceCommon in DevOps and SRE teams to ensure system reliability and uptime

The Observability Manager primarily oversees monitoring and logging strategies, ensuring system visibility, while the Site Reliability Engineer is more hands-on, focusing on automating infrastructure and maintaining system reliability. Both roles require technical expertise and often collaborate closely but differ in scope and daily responsibilities.

What are popular job titles related to Observability Manager jobs in Tennessee?

For Observability Manager jobs in Tennessee, the most frequently searched job titles are:

What job categories do people searching Observability Manager jobs in Tennessee look for?

The top searched job categories for Observability Manager jobs in Tennessee are:

What cities in Tennessee are hiring for Observability Manager jobs?

Cities in Tennessee with the most Observability Manager job openings:

Infographic showing various Observability Manager job openings in Tennessee as of June 2026, with employment types broken down into 100% Full Time. Highlights an 77% Physical, 5% Hybrid, and 18% Remote job distribution.

Infrastructure Services Engineer (Hybrid Eligible)

Oak Ridge, TN • On-site


Oak Ridge National Laboratory
Scientific Research and Development Services • 5 - 10K employees

8.8

Company rating: 8.8 out of 10

Based on 16 frontline employees who took The Breakroom Quiz

14th of 120 rated laboratories

Great coworkers

People enjoy working here

Good employer


$95 - $130/hr

Other

Medical, Dental, Vision, Life, Retirement, PTO

Posted 5 days ago


Job description

Select how often (in days) to receive an alert:

We are seeking an Infrastructure Services Engineer who will focus on specializing in monitoring and observability. This position resides in the Infrastructure Operations Center (IOC) in the Digital Services Infrastructure & Operations division of the Information Technology Services Directorate, atOak Ridge National Laboratory (ORNL).

As part of our team, you will design, operate, and continuously improve monitoring solutions across on-premises, cloud, and containerized environments. The IOC provides 24/7/365 monitoring and operational support for ORNL’s enterprise infrastructure and business-essential systems and services.

Major Duties/Responsibilities:
  • Design, implement, administer, and maintain enterprise monitoring and observability solutions across on-premises, cloud, and containerized environments.
  • Develop and optimize alerts, dashboards, reports, synthetic monitors, and telemetry pipelines to identify degradation early and accelerate incident triage and root-cause analysis.
  • Evaluate monitoring coverage, gaps, overlaps, and underused capabilities, and implement tools, integrations, and data sources that improve operational visibility.
  • Automate monitoring deployment, configuration, data collection, and remediation using PowerShell, Python, or other scripting and automation tools.
  • Evaluate and apply AI-assisted capabilities for anomaly detection, predictive analytics, and operational efficiency.
  • Collaborate with infrastructure, network, application, security, and other technical teams to improve system health and observability.
  • Support incident and problem management by providing relevant metrics, logs, performance trends, and historical analysis.
  • Work with vendors and internal subject matter experts to troubleshoot monitoring agents, collectors, integrations, and platform components.
  • Establish and maintain monitoring standards, topology diagrams, technical documentation, runbooks, and team procedures.
  • Support patching, backup, upgrade, and lifecycle activities for monitoring platforms and related infrastructure components.
  • Continuously improve alert thresholds, dashboards, data quality, automated remediations, and monitoring workflows to reduce noise and repetitive operational work.
  • Provide escalated support for monitoring-related issues and participate in an on-call or planned maintenance rotation as required.
  • Deliver ORNL’s mission by aligning behaviors, priorities, and interactions with our core values of Impact, Integrity, Teamwork, Safety, and Service. Promote equal opportunity by fostering a respectful workplace – in how we treat one another, work together, and measure success.
Basic Qualifications:
  • BS degree in information technology or a related technical field and 2 years of relevant experience.
  • Experience operating, administering, or engineering enterprise monitoring platforms for infrastructure, applications, networks, or cloud environments.
  • Experience supporting enterprise Windows and Linux server environments, including performance analysis and troubleshooting.
  • Experience developing automated solutions using PowerShell, Python, or similar scripting tools.
  • Working knowledge of cloud infrastructure, container platforms, orchestration technologies, virtualization, and virtual-machine lifecycle operations.
  • Understanding of networking fundamentals, system performance indicators, telemetry, and diagnostic methodologies.
Preferred Qualifications:
  • Strong analytical and problem-solving skills, including the ability to use operational data to identify issues and recommend improvements.
  • Strong written and verbal communication, customer service, collaboration, and technical documentation skills.
  • Ability to prioritize responsibilities and balance project work, operational support, and incident response in a fast-paced environment.
  • Demonstrated experience automating repetitive work or improving technical and operational processes.
  • Experience engineering and administering one or more enterprise-scale monitoring platforms, such as Prometheus, Grafana, Elastic, SolarWinds, or Dynatrace.
  • Experience with observability concepts and technologies, including metrics, logs, traces, baselining, synthetic monitoring, and service-level objectives.
  • Experience with anomaly detection, predictive analytics, AIOps, or automated remediation.
  • Knowledge of automation and infrastructure-as-code frameworks, such as Ansible, Terraform, or Azure Automation.
  • Experience using version-control systems to maintain scripts, configurations, dashboards, or infrastructure code.
  • Experience with virtualized or clustered compute environments, including performance tuning and lifecycle automation.
  • Familiarity with enterprise storage technologies, including direct-attached, SAN, and object storage, and their monitoring requirements.
  • Knowledge of enterprise server, storage, network hardware, and platform-level instrumentation.
  • Experience with enterprise backup, patching, configuration, or lifecycle management practices.
  • Understanding of change management and controlled operational workflows.
  • Experience working in regulated, scientific, government, or similarly complex technical environments.
  • Motivated self-starter with the ability to work independently and participate creatively in collaborative teams across the laboratory.
Special Requirements:
  • Visa sponsorship: Visa sponsorship is not available for this position.
  • Security, Credentialing, and Eligibility Requirements: Q Clearance: This position requires the ability to obtain and maintain a clearance from the Department of Energy. As such, this position is a Workplace Substance Abuse (WSAP) testing designated position. WSAP positions require passing a pre-placement drug test and participation in an ongoing random drug testing program.
  • For Hybrid eligible positions: In addition, we offer a flexible work environment that supports both the organization and the employee. A hybrid/onsite working arrangement may be available with this position.
About ORNL:

As a U.S. Department of Energy (DOE) Office of Science national laboratory, ORNL has an impressive 80-year legacy of addressing the nation’s most pressing challenges. Our team is made up of over 7,000 dedicated and innovative individuals! Our goal is to create an environment where a variety of perspectives and backgrounds are valued, ensuring ORNL is known as a top choice for employment. These principles are essential for supporting our broader mission to drive scientific breakthroughs and translate them into solutions for energy, environmental, and security challenges facing the nation.

ORNL offers competitive pay and benefits programs to attract and retain individuals who demonstrate exceptional work behaviors. The laboratory provides a range of employee benefits, including medical and retirement plans and flexible work hours, to support the well-being of you and your family. Employee amenities such as on-site fitness, banking, and cafeteria facilities are also available for added convenience.

Other benefits include the following: Prescription Drug Plan, Dental Plan, Vision Plan, 401(k) Retirement Plan, Contributory Pension Plan, Life Insurance, Disability Benefits, Generous Vacation and Holidays, Parental Leave, Legal Insurance with Identity Theft Protection, Employee Assistance Plan, Flexible Spending Accounts, Health Savings Accounts, Wellness Programs, Educational Assistance, Relocation Assistance, and Employee Discounts.

This position will remain open for a minimum of 5 days after which it will close when a qualified candidate is identified and/or hired.

ORNL is an equal opportunity employer. All qualified applicants, including individuals with disabilities and protected veterans, are encouraged to apply. UT-Battelle is an E-Verify employer.

#J-18808-Ljbffr


What Oak Ridge National Laboratory employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom