1

Kafka Sre Jobs (NOW HIRING)

Senior Kafka SRE Engineer

Austin, TX · On-site

$120K - $155K/yr

We are seeking a Kafka Site Reliability Engineer to help build, operate, and continuously improve Schwab's enterprise streaming platform ecosystem. This role combines deep expertise in Confluent ...

Senior Kafka SRE Engineer

Austin, TX · On-site

$56.50 - $75/hr

We are seeking a Kafka Site Reliability Engineer to help build, operate, and continuously improve Schwab's enterprise streaming platform ecosystem. This role combines deep expertise in Confluent ...

Our SRE team combines software engineering, systems engineering, and Devops practices to build and ... You will help build next-generation Kafka infrastructure and platform services, collaborating cross ...

Site Reliability Engineer

Chicago, IL · On-site

$58.75 - $78/hr

... Kafka, and Kubernetes shop. What You Will Do * Technical Vision & Roadmap: Define the 12-18 month ... Level up the entire SRE organization through design reviews, architectural "office hours, " and ...

Site Reliability Engineer (SRE)

Austin, TX · On-site

$56.50 - $75/hr

Site Reliability Engineer (SRE) Location: Austin, TX Job Type: Full Time Technical Skills: * 6+ ... Kafka, Cloud SQL, etc. * Strong experience in using industry standard monitoring tools e.g ...

SRE

$58.25 - $77.50/hr

These components include Kubernetes workloads (MuleSoft, Java) and Kafka. Together with two ... SRE maturity and reliability of the applications.

Site Reliability Engineer

Manhattan, NY · On-site

$62.75 - $83.50/hr

Site Reliability Engineer (SRE) Equity Trading Platform Location: New York, NY Experience: 5+ Years ... IBM MQ, Kafka, Tibco EMS, ActiveMQ, or AWS SQS/SNS . * Jenkins, Git, Terraform, Ansible, Jira ...

Site Reliability Engineer

Palo Alto, CA · On-site

$67 - $89/hr

About the DevOps / SRE Team The DevOps/SRE team at Obsidian ensures that engineering excellence ... Bonus if you're familiar with technologies like Kafka, Elasticsearch, PostgreSQL, ScyllaDB ...

Site Reliability Engineer (SRE)

Austin, TX · On-site

$56.50 - $75/hr

Site Reliability Engineer (SRE) Location: Austin, TX Job Type: Full Time Job Summary - Seasoned ... Kafka etc. o Automate of day-day operational tasks. o Be part of the Exit reviews to ensure the ...

Staff Site Reliability Engineer (SRE) - Platform Engineering Note: This position follows a hybrid ... Deep-seated expertise in GCP (Networking, IAM, GKE) and the ability to scale Kafka clusters for ...

Staff Site Reliability Engineer (SRE) - Platform Engineering Note: This position follows a hybrid ... Kafka, and Kubernetes shop. What You Will Do * Technical Vision & Roadmap: Define the 12-18 month ...

Site Reliability Engineer (SRE)

Omaha, NE · On-site

$54.50 - $72.50/hr

Site Reliability Engineer (SRE) Location: Omaha, NE / Dallas, TX Job Type: Full Time Job Summary ... Kafka etc. o Automate of day-day operational tasks. o Be part of the Exit reviews to ensure the ...

Staff Site Reliability Engineer (SRE) - Platform Engineering Note: This position follows a hybrid ... Deep-seated expertise in GCP (Networking, IAM, GKE) and the ability to scale Kafka clusters for ...

Site Reliability Engineer (SRE)

Plano, TX · On-site

$54.75 - $72.75/hr

... Kafka,Bash,CICD,databases,automation pipelines,DevOps environment,Docker Kubernetes,Grafana,Jenkins,Azure Google Cloud Platform,monitoring tools,Prometheus,Python,system reliability,SRE practices ...

$57.75 - $76.75/hr

Site Reliability Engineer (SRE) Department: Technology Location: Manila Reporting To: Head of Infra ... Exposure to cutting-edge AI and big data infrastructure (Spark, Kafka, ScyllaDB, Flink). Employment ...

Site Reliability Engineer

Riverwoods, IL · On-site

$59.25 - $78.75/hr

We are looking for a Site Reliability Engineer to join our growing team in Riverwoods, IL United ... Expertise in Message Broker (preferably Rabbit MQ, Kafka). * Expertise on Hadoop, spark commands ...

next page

Showing results 1-20

Kafka Sre information

See salary details

$10

$63

$91

How much do kafka sre jobs pay per hour?

As of Sep 10, 2026, the average hourly pay for kafka sre in the United States is $63.74, according to ZipRecruiter salary data. Most workers in this role earn between $54.81 and $72.84 per hour, depending on experience, location, and employer.

What are popular job titles related to Kafka Sre jobs?

For Kafka Sre jobs, the most frequently searched job titles are:

Senior Kafka SRE Engineer

Austin, TX • On-site

$120K - $155K/yr

Full-time

Re-posted 5 days ago


Key responsibilities

  • Build, operate, and improve Schwab's enterprise streaming platform ecosystem using Confluent Kafka technologies.

  • Design, deploy, and support Kafka environments across on-premises and cloud platforms, ensuring high availability and resilience.

  • Drive operational excellence through automation, observability, infrastructure as code, and AIOps-driven capabilities.


Job description

Your Opportunity
Your Opportunity
At Schwab, you're empowered to make an impact on your career. Here, innovative thought meets creative problem solving, helping us challenge the status quo and transform the finance industry together. We believe in the importance of in-office collaboration and fully intend for the selected candidate for this role to work on site in the specified location(s).
We are seeking a Kafka Site Reliability Engineer to help build, operate, and continuously improve Schwab's enterprise streaming platform ecosystem. This role combines deep expertise in Confluent Kafka technologies with modern Site Reliability Engineering practices to deliver highly available, secure, and resilient streaming services that support critical business capabilities across the firm.
As a member of the team, you will drive operational excellence through automation, observability, Infrastructure as Code, and AIOps-driven capabilities. You will play a key role in designing, deploying, and supporting Confluent Kafka environments across on-premises and cloud platforms while helping engineering teams deliver reliable real-time data solutions. Success in this role requires the ability to proactively identify risks, solve complex technical challenges, improve platform performance, and enhance system reliability through data-driven decision-making and continuous improvement.
You will contribute to the evolution of Kafka infrastructure supporting capabilities such as Schema Registry, Kafka Connect, ksqlDB, Cluster Linking, and cloud-native deployments. Through the application of automation, predictive analytics, and AI-driven operational insights, you will help reduce operational toil, strengthen platform resiliency, improve incident response, and accelerate issue resolution.
The ideal candidate thrives in highly distributed environments and enjoys partnering with engineers, architects, and infrastructure teams to improve platform health, streamline deployment processes, increase observability, and establish best practices across the streaming ecosystem. This role offers the opportunity to influence the future of event streaming at Schwab while developing expertise in emerging AIOps, cloud, automation, and reliability engineering capabilities.
What you have
Required Qualifications
  • 5-7 years of experience supporting and administering enterprise-scale production Confluent Kafka platforms, including Confluent Platform and Confluent Cloud.
  • 5-7 years of experience developing Python automation, operational tooling, observability dashboards, and alerting solutions.
  • Hands-on experience applying AIOps concepts, including anomaly detection, event correlation, alert reduction, predictive analytics, automated remediation, and AI-driven operational insights within production environments.
  • Experience leveraging predictive monitoring and intelligent automation to improve platform reliability, incident management, and operational efficiency.
  • Hands-on experience deploying and operating Kafka platforms on Kubernetes, preferably Google Kubernetes Engine (GKE), using Helm charts and cloud-native operational practices.
  • Strong experience using Terraform and Ansible to automate infrastructure provisioning, configuration management, and platform lifecycle activities.
  • Deep expertise with the Confluent Kafka ecosystem, including ZooKeeper, KRaft, Schema Registry, Kafka Connect, ksqlDB, Cluster Linking, MirrorMaker, and role-based access controls (RBAC).
  • 3+ years of experience working with public cloud technologies, with Google Cloud Platform (GCP) preferred.
  • Strong understanding of Kafka architecture and client internals, including producers, consumers, partitions, replication, serialization, consumer groups, performance tuning, and exactly-once processing.
  • Strong Linux administration, troubleshooting, performance tuning, and networking experience supporting high-throughput distributed systems.
  • Experience supporting large-scale distributed systems, highly available platforms, fault-tolerant architectures, and production-critical workloads.
  • Experience performing incident response, root-cause analysis, problem resolution, and post-incident continuous improvement activities.
  • Experience implementing and maintaining observability solutions, operational metrics, service-level objectives (SLOs), dashboards, and actionable monitoring controls.
  • Bachelor's degree in Computer Science, Engineering, Information Technology, or a related discipline.

Preferred Qualifications
  • Confluent certifications such as CCDAK or CCAAK.
  • Google Cloud certifications such as Professional Cloud Architect or Professional Cloud DevOps Engineer.
  • Experience operating Confluent for Kubernetes (CFK) and Confluent Cloud APIs.
  • Experience implementing or supporting AIOps platforms for predictive incident detection, automated root-cause analysis, and self-healing infrastructure capabilities.
  • Experience with observability and monitoring technologies such as Grafana, InfluxDB, BigQuery, Prometheus, Splunk, Datadog, or comparable monitoring platforms.
  • Experience with CI/CD technologies such as GitHub Actions, Cloud Build, or similar automation frameworks.
  • Experience supporting additional messaging and streaming technologies such as RabbitMQ, IBM MQ, Solace, or Google Pub/Sub.
  • Understanding of modern Site Reliability Engineering practices, including SLIs, SLOs, error budgets, reliability engineering, and operational excellence methodologies.
  • Strong communication, collaboration, and relationship-building skills with the ability to effectively partner across technical and business teams.
  • Demonstrated ability to adapt to changing priorities, drive initiatives independently, and maintain a strong sense of ownership and accountability.

In addition to the salary range, this role is eligible for bonus or incentive opportunities.