1

Observability Kubernetes Jobs (NOW HIRING)

Kubernetes Engineer

Phoenix, AZ · On-site

$56.50 - $75.25/hr

Implement observability stacks using tools such as Prometheus, Grafana, Alertmanager, Splunk, ELK ... Apply expertise in Kubernetes networking concepts, including Ingress, Services, CNI plugins, and ...

The role requires a strong understanding of observability tools and practices, with a focus on Prometheus, Grafana, Gardener Kubernetes, and Splunk. Experience with Dynatrace is a plus. Skills: • ...

Observability Engineer

Phoenix, AZ · On-site

$52.50 - $71.75/hr

Observability Engineer (Dynatrace, Splunk & OpenSearch) Location: Phoenix, AZ (Onsite) Long Term ... Implement monitoring for cloud-native applications, containers, Kubernetes, and microservices.

The ideal candidate will have deep expertise in Dynatrace, Splunk, OpenSearch/Elasticsearch, Kubernetes, Linux, and cloud-native observability solutions. Experience leveraging AI/ML and Generative AI ...

Kubernetes Engineer

Charlotte, NC · On-site

$51.50 - $70.50/hr

Azure Kubernetes Service (AKS) Management * Design, deploy, and manage enterprise-scale AKS ... Configure traffic splitting, load balancing, and observability with Linkerd * Troubleshoot service ...

Senior Observability Engineer

Natick, MA · On-site

$122K - $189K/yr

Build scalable, multi-tenant observability solutions for Kubernetes clusters running microservices at scale. * Implement SLOs, SLIs, and error budgets-integrating observability into SRE practices.

Kubernetes Engineer

Charlotte, NC · On-site

$51.50 - $70.50/hr

Azure Kubernetes Service (AKS) Management * Design, deploy, and manage enterprise-scale AKS ... Configure traffic splitting, load balancing, and observability with Linkerd * Troubleshoot service ...

Kubernetes Engineer

Charlotte, NC · On-site

$51.50 - $70.50/hr

Azure Kubernetes Service (AKS) Management * Design, deploy, and manage enterprise-scale AKS ... Configure traffic splitting, load balancing, and observability with Linkerd * Troubleshoot service ...

Kubernetes Architect

Herndon, VA · On-site

$60 - $70/hr

Kubernetes Architect Location-Type: Hybrid (Herndon, VA, Austin, TX, or Newtown Square, PA ... Implement secure monitoring and observability practices across cloud environments * Present ...

Kubernetes Architect

Herndon, VA · Remote

$60 - $75/hr

Kubernetes Architect, location is Hybrid/Remote. The start date is July 2026 for this 12-month ... Support monitoring, observability, and signal management initiatives * Establish platform best ...

C. is seeking a Kubernetes Administrator to manage and optimize production clusters across Amazon ... The role involves overseeing cluster operations, implementing observability stacks, and building CI ...

Leidos is seeking a Principal Kubernetes Engineer to support the design, development, and ... Improve platform observability through metrics, logging, tracing, dashboards, and proactive ...

next page

Showing results 1-20

Observability Kubernetes information

What cities are hiring for Observability Kubernetes jobs? Cities with the most Observability Kubernetes job openings:
What states have the most Observability Kubernetes jobs? States with the most job openings for Observability Kubernetes jobs include:
What job categories do people searching Observability Kubernetes jobs look for? The top searched job categories for Observability Kubernetes jobs are:

Lead Site Reliability Engineer, Factory Software

Tesla

Fremont, CA • On-site

$62.50 - $83.25/hr

Full-time

Re-posted 9 days ago


Tesla rating

8.5

Company rating: 8.5 out of 10

Based on 681 frontline employees who took The Breakroom Quiz

1st of 44 rated automakers


Job description

Job Summary:
Tesla is building critical applications to enable manufacturing and warehouse management with a strong emphasis on reliability, availability, scalability, speed, and security. As the Lead Site Reliability Engineer, you will be the primary technical owner and leader for the Factory Software team’s reliability, observability, and infrastructure strategy, combining deep hands-on engineering with leadership to ensure the full stack is highly reliable and performant.
Responsibilities:
• Provide technical leadership and set the vision for observability, reliability, and platform standardization across the Factory Software team
• Design and implement end-to-end observability and telemetry solutions (OTEL, Prometheus, Grafana, Tempo, etc.) while mentoring the team on best practices
• Own the reliability of the full stack: Kubernetes infrastructure, virtual machines, databases, and the middleware applications connecting PLCs, MES systems, and other factory services
• Define and drive SLIs, SLOs, error budgets, and golden signals across services
• Lead major initiatives to eliminate speed bottlenecks, database contention, and infrastructure issues through proactive monitoring and automation
• Write production-grade code and build tools to reduce toil and improve deployment, monitoring, and operational workflows
• Participate hands-on in on-call rotations, live troubleshooting during outages (NOC bridges), and blameless post-mortems
• Collaborate closely with Platform Engineering, Infrastructure, Controls Engineering, and Software Engineering teams to embed reliability and observability into architecture and development practices
• Mentor and coach engineers on technical excellence, observability, Kubernetes, Linux, networking, and reliable system design
• Drive continuous improvement in incident response, system performance, and engineering standards across the team
Qualifications:
Required:
• 7+ years of experience in Site Reliability Engineering, Platform Engineering, or related systems roles, with significant hands-on experience at scale
• Strong technical expertise in Kubernetes, Docker, Linux administration, and networking (routing, VLANs, firewalls, load balancers)
• Deep experience with observability tools and concepts (Prometheus, Grafana, Tempo, OTEL, Splunk, etc.)
• Proven track record of designing and implementing reliable, observable distributed systems
• Proficiency in at least one high-level language (Go, Python, or Java) with experience writing production-grade code
• Demonstrated ability to lead technical initiatives and raise the engineering bar without formal people management authority
• Experience with on-call rotations, incident command, and driving reliability improvements through blameless post-mortems
• Strong bias for action, excellent communication skills, and a desire to mentor and uplift other engineers
Preferred:
• Experience in manufacturing, industrial automation, or complex operational environments is a strong plus
Company:
Tesla is an electric vehicle and clean energy company that provides electric cars, solar, and renewable energy solutions. Founded in 2003, the company is headquartered in Austin, USA, with a team of 10001+ employees. The company is currently Late Stage.

What Tesla employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom