1

Observability Manager Jobs in Tennessee (NOW HIRING)

Lead Principal Site Reliability Engineer

Nashville, TN · On-site

$55 - $73.25/hr

Proven expertise with infrastructure-as-code and configuration-management technologies such as Terraform, Ansible, Chef, Puppet, or equivalent tools. * Experience implementing observability solutions ...

OpenShift (OCP) Observability & Troubleshooting Tooling Experience using one or more of the ... Leadership & Management Skills At least 5 years' experience leading teams with increasing levels of ...

DevOps Lead

Brentwood, TN · On-site

$120K - $140K/yr

Implement observability, SLI/SLO/SLA dashboards, production incident response, RCA, and problem management. * Integrate DevSecOps controls including security scanning, vulnerability management, Key ...

Senior Platform Software Engineer

Nashville, TN · On-site

$118K - $156K/yr

Familiarity with Kubernetes internals, controllers/operators, cluster lifecycle management, networking, or storage. * Experience with observability platforms, including logging, metrics, tracing, and ...

Sr Product Manager, Technical Leader

Nashville, TN · On-site

$161K - $186K/yr

... observability needs for modern cloud native infrastructure. The flagship technologies, Cilium and ... Your Impact As the Sr Product Manager, Technical Leader , you are the voice of the customer ...

New

Partner closely with Security Engineering, Product Management, SRE, and OCI platform teams to define and deliver next-generation application security capabilities. * Establish robust observability ...

... Observability, Finance, Business Operations, hardware partners, and senior leadership to deliver ... As a Principal Program Manager , you will oversee the planning, execution, and operational ...

... Observability, Finance, Business Operations, hardware partners, and senior leadership to deliver ... As a Principal Program Manager , you will oversee the planning, execution, and operational ...

Senior Platform Software Engineer

Nashville, TN · On-site

$118K - $156K/yr

Familiarity with Kubernetes internals, controllers/operators, cluster lifecycle management, networking, or storage. * Experience with observability platforms, including logging, metrics, tracing, and ...

Partner closely with Security Engineering, Product Management, SRE, and OCI platform teams to define and deliver next-generation application security capabilities. * Establish robust observability ...

Principal AI Agent / ML Software Engineer

Nashville, TN · On-site

$130K - $174K/yr

Define service boundaries, APIs, data models, state management, consistency tradeoffs, failure ... observability, and performance analysis. * Experience defining SLIs/SLOs, production readiness ...

New

... Observability, Finance, Business Operations, hardware partners, and senior leadership to deliver ... As a Principal Program Manager , you will oversee the planning, execution, and operational ...

Showing results 41-60

Observability Manager information

What is the difference between Observability Manager vs Site Reliability Engineer?

AspectObservability ManagerSite Reliability Engineer
CredentialsTypically requires experience in monitoring, logging, and cloud tools; certifications like AWS, Google Cloud, or Kubernetes are commonRequires strong background in systems engineering, scripting, and cloud platforms; certifications like AWS, GCP, or Linux are often preferred
Work EnvironmentFocuses on overseeing observability tools, data analysis, and team coordination in tech environmentsHands-on role involving system automation, incident response, and infrastructure reliability
Industry UsageUsed across tech companies to improve system visibility and performanceCommon in DevOps and SRE teams to ensure system reliability and uptime

The Observability Manager primarily oversees monitoring and logging strategies, ensuring system visibility, while the Site Reliability Engineer is more hands-on, focusing on automating infrastructure and maintaining system reliability. Both roles require technical expertise and often collaborate closely but differ in scope and daily responsibilities.

What are the most commonly searched types of Observability jobs in Tennessee? The most popular types of Observability jobs in Tennessee are:
What are popular job titles related to Observability Manager jobs in Tennessee? For Observability Manager jobs in Tennessee, the most frequently searched job titles are:
What job categories do people searching Observability Manager jobs in Tennessee look for? The top searched job categories for Observability Manager jobs in Tennessee are:
Infographic showing various Observability Manager job openings in Tennessee as of June 2026, with employment types broken down into 100% Full Time. Highlights an 77% Physical, 5% Hybrid, and 18% Remote job distribution.

Lead Principal Site Reliability Engineer

Oracle

Nashville, TN • On-site

$55 - $73.25/hr

Full-time

Medical, Dental, Vision, Life, Retirement, PTO

Posted 23 days ago


Oracle rating

8.7

Company rating: 8.7 out of 10

Based on 151 frontline employees who took The Breakroom Quiz

55th of 242 rated software companies


Job description


Serves as a consultant and leads the design and architecture of infrastructure and service, ensuring alignment with reliability and functionality standards. Takes full ownership of forecasting of demands and responding to capacity needs. Owns collaborations with software development teams to develop reliable and scalable infrastructures. Recommends methods for performing data collection to maintain and optimize operations and reliability. Oversees incident response and/or maintenance tasks. Provides strategic, future-oriented health and performance reporting. Contributes to strategies for automation and reviews the development and implementation of automation. Provides expert-level communication about services and anticipates, analyzes, and explains the impact of changes, considering strategic goals. Serves as a role model in providing support for technology and reviews documentation for accuracy. Leads the implementation of innovative tools and provides expertise in site reliability trends.
Responsibilities
Key Responsibilities
Reliability Strategy and Technical Leadership
  • Define and drive the site reliability engineering strategy for large-scale, distributed, and business-critical platforms.
  • Establish reliability standards, engineering practices, and operational readiness requirements across multiple teams.
  • Serve as a senior technical authority for system reliability, scalability, resilience, performance, and production operations.
  • Influence architecture and design decisions to ensure systems are supportable, observable, fault tolerant, and capable of meeting availability objectives.
  • Identify systemic reliability risks and lead cross-functional initiatives to address them.
  • Provide technical direction and mentorship to site reliability engineers, software engineers, platform engineers, and operations teams.
  • Lead technical reviews and promote consistent engineering practices across the organization.

Service Reliability and Observability
  • Define and implement service-level indicators, service-level objectives, error budgets, and operational health metrics.
  • Develop comprehensive monitoring, logging, tracing, alerting, and observability strategies.
  • Improve the quality and actionability of alerts while reducing unnecessary operational noise.
  • Establish dashboards and reporting mechanisms that clearly communicate service health, performance, capacity, and risk.
  • Use production data and reliability trends to prioritize engineering investments and continuous-improvement initiatives.

Automation and Platform Engineering
  • Design and implement automation that reduces manual intervention, operational toil, and human error.
  • Build or enhance tools for deployment, configuration management, infrastructure provisioning, incident response, and service recovery.
  • Promote infrastructure-as-code, policy-as-code, automated testing, and repeatable deployment practices.
  • Partner with development teams to improve continuous integration and continuous delivery pipelines.
  • Develop self-healing and automated remediation capabilities where appropriate.
  • Contribute production-quality software and reusable platform components using modern programming and scripting languages.

Incident Management and Problem Resolution
  • Provide technical leadership during complex, high-severity production incidents.
  • Coordinate diagnosis, containment, recovery, and stakeholder communication during service disruptions.
  • Lead blameless post-incident reviews and ensure that corrective actions address root causes rather than symptoms.
  • Identify recurring failure patterns and develop long-term engineering solutions.
  • Improve incident-management processes, escalation procedures, runbooks, and recovery playbooks.
  • Participate in an on-call rotation or provide senior escalation support for critical services, as required.

Capacity, Performance, and Resilience
  • Lead capacity planning, performance analysis, load testing, and demand forecasting for critical platforms.
  • Identify performance bottlenecks and recommend architectural or operational improvements.
  • Design and validate high-availability, disaster-recovery, backup, and business-continuity capabilities.
  • Lead resilience testing, failure-mode analysis, game days, and controlled fault-injection exercises.
  • Ensure recovery-time and recovery-point objectives are defined, tested, and achievable.

Security and Operational Governance
  • Partner with security and compliance teams to embed security into infrastructure, automation, and operational practices.
  • Support vulnerability remediation, access-control improvements, audit readiness, and secure configuration management.
  • Ensure production environments meet organizational standards for change management, data protection, and operational governance.
  • Balance reliability, security, delivery speed, cost, and business priorities when recommending technical solutions.

Required Qualifications
  • Extensive professional experience in site reliability engineering, software engineering, cloud infrastructure, platform engineering, systems engineering, or a related technical discipline.
  • Demonstrated experience designing, operating, and improving highly available production systems at significant scale.
  • Deep knowledge of distributed systems, cloud architecture, networking, operating systems, storage, databases, and service dependencies.
  • Advanced experience with at least one major cloud platform, such as Oracle Cloud Infrastructure, Amazon Web Services, Microsoft Azure, or Google Cloud Platform.
  • Strong experience with containerization and orchestration technologies, including Docker and Kubernetes.
  • Proven expertise with infrastructure-as-code and configuration-management technologies such as Terraform, Ansible, Chef, Puppet, or equivalent tools.
  • Experience implementing observability solutions using metrics, logs, traces, dashboards, and automated alerting.
  • Strong programming or scripting skills in one or more languages such as Python, Go, Java, JavaScript, Bash, or similar.
  • Experience with continuous integration, continuous delivery, automated testing, and modern release-management practices.
  • Demonstrated leadership during critical production incidents and complex technical investigations.
  • Ability to diagnose difficult system issues across applications, infrastructure, networks, databases, and cloud services.
  • Strong written and verbal communication skills, including the ability to explain technical risk and recommendations to engineering leaders and business stakeholders.
  • Proven ability to lead cross-functional technical initiatives without relying solely on formal authority.

Preferred Qualifications
  • Experience supporting enterprise-scale cloud services, software-as-a-service platforms, or other high-availability customer-facing systems.
  • Experience defining and operating service-level objectives, error budgets, and reliability scorecards.
  • Knowledge of chaos engineering, resilience testing, and automated recovery techniques.
  • Experience with multi-region, hybrid-cloud, or multi-cloud architectures.
  • Familiarity with security frameworks, compliance requirements, and regulated operating environments.
  • Experience improving cloud cost efficiency, capacity utilization, or infrastructure performance.
  • Contributions to internal engineering standards, technical communities, open-source projects, or industry publications.
  • Bachelor's or advanced degree in computer science, engineering, information systems, or a related field, or equivalent practical experience.

Qualifications
Disclaimer:
Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.
Range and benefit information provided in this posting are specific to the stated locations only
US: Hiring Range in USD from: $96,300 to $264,100 per annum. May be eligible for bonus, equity, and compensation deferral.
Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.
Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.
Oracle US offers a comprehensive benefits package which includes the following:
1. Medical, dental, and vision insurance, including expert medical opinion
2. Short term disability and long term disability
3. Life insurance and AD&D
4. Supplemental life insurance (Employee/Spouse/Child)
5. Health care and dependent care Flexible Spending Accounts
6. Pre-tax commuter and parking benefits
7. 401(k) Savings and Investment Plan with company match
8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
9. 11 paid holidays
10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
11. Paid parental leave
12. Adoption assistance
13. Employee Stock Purchase Plan
14. Financial planning and group legal
15. Voluntary benefits including auto, homeowner and pet insurance
The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.
Career Level - IC5
About Us
Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.
True innovation starts when everyone is empowered to contribute. That's why we're committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.
We're committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing accommodation-request_mb@oracle.com or by calling 1-888-404-2494 in the United States.
Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans' status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

What Oracle employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom


Oracle logo

About Oracle

Sourced by ZipRecruiter

An Oracle career can span industries, roles, Countries and cultures, giving you the opportunity to flourish in new roles and innovate, while blending work life in. Oracle has thrived through 40+ years of change by innovating and operating with integrity while delivering for the top companies in almost every industry. In order to nurture the talent that makes this happen, we are committed to an inclusive culture that celebrates and values diverse insights and perspectives, a workforce that inspires thought leadership and innovation. Oracle offers a highly competitive suite of Employee Benefits designed on the principles of parity, consistency, and affordability. The overall package includes certain core elements such as Medical, Life Insurance, access to Retirement Planning, and much more. We also encourage our employees to engage in the culture of giving back to the communities where we live and do business. At Oracle, we believe that innovation starts with diversity and inclusion and to create the future we need talent from various backgrounds, perspectives, and abilities. We ensure that individuals with disabilities are provided reasonable accommodation to successfully participate in the job application, interview process, and in potential roles. to perform crucial job functions. That's why we're committed to creating a workforce where all individuals can do their best work. It's when everyone's voice is heard and valued that we're inspired to go beyond what's been done before.

Industry

It services

Company size

10,000+ Employees

Headquarters location

Redwood City, CA, US

Year founded

1977

Social media