1

Manager Clickhouse Jobs in Houston, TX (NOW HIRING)

Manager Clickhouse information

What is the difference between Manager Clickhouse vs Data Engineer?

AspectManager ClickhouseData Engineer
Primary FocusOversees Clickhouse database management, team coordination, and strategyDesigns, develops, and maintains data pipelines and infrastructure
Required SkillsDatabase administration, team leadership, project managementETL processes, SQL, programming, data modeling
CertificationsDatabase management, cloud certifications, project managementData engineering, cloud, and programming certifications
Work EnvironmentManagement, strategic planning, team collaborationTechnical development, coding, data architecture

The Manager Clickhouse primarily focuses on overseeing Clickhouse database operations and leading teams, while a Data Engineer concentrates on building and maintaining data pipelines and infrastructure. Both roles require technical knowledge, but the Manager Clickhouse emphasizes management skills, whereas the Data Engineer emphasizes technical development and data architecture.

What are popular job titles related to Manager Clickhouse jobs in Houston, TX?

For Manager Clickhouse jobs in Houston, TX, the most frequently searched job titles are:

What cities near Houston, TX are hiring for Manager Clickhouse jobs?

Cities near Houston, TX with the most Manager Clickhouse job openings:

Staff Observability Platform Engineer (SRE)

Recruitment.ai

Houston, TX • On-site

$54.50 - $72.25/hr

Other

Posted 5 days ago


Job description

Staff Observability Platform Engineer (SRE)
Locations: Seattle, WA (Hybrid), Houston, TX (Hybrid), New York, NY (Hybrid)
What You’ll Do
  • Design, build, and evolve observability platforms across metrics, logs, traces, alerting, and telemetry pipelines.
  • Lead the implementation of scalable observability solutions that support Nscale’s growing GPU and AI infrastructure.
  • Partner with SRE, infrastructure, platform, and AI/ML teams to ensure observability is embedded throughout the software and infrastructure lifecycle.
  • Drive improvements in monitoring coverage, alert quality, service health visibility, and incident response effectiveness.
  • Develop standards, frameworks, and reusable patterns that simplify observability adoption across engineering teams.
  • Identify reliability risks and operational blind spots, helping teams proactively address them before they impact customers.
  • Contribute to architectural decisions around telemetry collection, storage, retention, cardinality management, and performance optimization.
  • Lead technical initiatives and projects that improve platform scalability, reliability, and operational efficiency.
  • Mentor engineers and provide technical guidance through design reviews, code reviews, and knowledge sharing.
  • Participate in incident investigations and postmortems, translating operational learnings into durable platform improvements.
  • Evaluate new observability technologies and practices, balancing innovation with operational simplicity and long-term maintainability.
About You
  • Experience in SRE, platform engineering, infrastructure engineering, observability engineering, or related disciplines.
  • Strong experience building and operating observability platforms in cloud-native, distributed environments.
  • Deep hands-on experience with several of the following technologies: Prometheus, Thanos, VictoriaMetrics, Grafana, Loki, Tempo, OpenTelemetry, ClickHouse, Elastic, or similar platforms.
  • Strong software engineering skills with proficiency in Go, Python, or equivalent languages.
  • Experience operating and troubleshooting Kubernetes-based platforms at scale.
  • Strong understanding of monitoring, logging, tracing, telemetry pipelines, and modern observability practices.
  • Experience designing systems with scalability, reliability, performance, and operational simplicity in mind.
  • Proficiency with Infrastructure-as-Code tools such as Terraform, Ansible, or equivalent.
  • Ability to lead technical initiatives and influence engineering decisions across multiple teams.
  • Excellent communication skills with the ability to explain technical tradeoffs and align stakeholders around pragmatic solutions.
Preferred
  • Experience operating observability systems in GPU, AI/ML, HPC, or large-scale compute environments.
  • Familiarity with Slurm, Kubernetes GPU scheduling, or AI infrastructure platforms.
  • Experience with high-volume telemetry pipelines and streaming technologies such as Kafka, Vector, or Fluent Bit.
  • Knowledge of observability challenges related to model training, inference workloads, GPU utilization, and distributed AI systems.
  • Experience mentoring engineers and helping grow technical capability across teams.