NVIDIA is seeking an Engineering Manager to lead the development of an agentic platform for ... Experience with OpenTelemetry, Prometheus, Grafana, Jaeger, ClickHouse, Elastic, or similar ...
NVIDIA is seeking an Engineering Manager to lead the development of an agentic platform for ... Experience with OpenTelemetry, Prometheus, Grafana, Jaeger, ClickHouse, Elastic, or similar ...
Sr. Database Reliability Engineer
Hawthorne, CA · On-site
$155 - $205/hr
Ownership of the health, uptime and performance of all PostgreSQL/Oracle/MySQL & Clickhouse ... Knowledge of configuration management automation and limiting the amount of manual toil * Knowledge ...
Sr. Database Reliability Engineer
Hawthorne, CA · On-site
$155 - $205/hr
Ownership of the health, uptime and performance of all PostgreSQL/Oracle/MySQL & Clickhouse ... Knowledge of configuration management automation and limiting the amount of manual toil * Knowledge ...
NVIDIA is seeking an Engineering Manager to lead the development of an agentic platform for ... Experience with OpenTelemetry, Prometheus, Grafana, Jaeger, ClickHouse, Elastic, or similar ...
NVIDIA is seeking an Engineering Manager to lead the development of an agentic platform for ... Experience with OpenTelemetry, Prometheus, Grafana, Jaeger, ClickHouse, Elastic, or similar ...
Senior Staff Software Engineer with Observability Platform
Redwood City, CA · On-site
$149K - $197K/yr
... and stream management). • Visualization: Expert-level proficiency in Grafana for creating ... Proven experience with Prometheus, Thanos, or ClickHouse and working within a structured Agile ...
Senior Staff Software Engineer with Observability Platform
Redwood City, CA · On-site
$149K - $197K/yr
... and stream management). • Visualization: Expert-level proficiency in Grafana for creating ... Proven experience with Prometheus, Thanos, or ClickHouse and working within a structured Agile ...
Senior Software Engineer, AI Data Systems & Database Infrastructure
Redwood City, CA · On-site
$168K - $205K/yr
Experience with analytical data stores such as ClickHouse, BigQuery, Snowflake, Redshift, Druid ... Experience managing cost and performance tradeoffs across cloud-managed and self-hosted database ...
Senior Software Engineer, AI Data Systems & Database Infrastructure
Redwood City, CA · On-site
$168K - $205K/yr
Experience with analytical data stores such as ClickHouse, BigQuery, Snowflake, Redshift, Druid ... Experience managing cost and performance tradeoffs across cloud-managed and self-hosted database ...
Sr. Database Reliability Engineer
Hawthorne, CA · On-site
$155K - $205K/yr
Ownership of the health, uptime and performance of all PostgreSQL/Oracle/MySQL & Clickhouse ... Knowledge of configuration management automation and limiting the amount of manual toil * Knowledge ...
Sr. Database Reliability Engineer
Hawthorne, CA · On-site
$155K - $205K/yr
Ownership of the health, uptime and performance of all PostgreSQL/Oracle/MySQL & Clickhouse ... Knowledge of configuration management automation and limiting the amount of manual toil * Knowledge ...
Business Intelligence Analyst
Santa Clara, CA · On-site
$140K - $193K/yr
... Kafka/ClickHouse stack. * Validate and reconcile data from disparate systems (MES, tool logs, recipe management, metrology platforms, yield databases) to ensure accurate linkage at the wafer, lot ...
Business Intelligence Analyst
Santa Clara, CA · On-site
$140K - $193K/yr
... Kafka/ClickHouse stack. * Validate and reconcile data from disparate systems (MES, tool logs, recipe management, metrology platforms, yield databases) to ensure accurate linkage at the wafer, lot ...
Sr. Database Reliability Engineer
Hawthorne, CA · On-site
$155 - $205/hr
Ownership of the health, uptime and performance of all PostgreSQL, Oracle, MySQL and ClickHouse ... Knowledge of configuration management automation and reducing manual toil * Knowledge of data ...
Sr. Database Reliability Engineer
Hawthorne, CA · On-site
$155 - $205/hr
Ownership of the health, uptime and performance of all PostgreSQL, Oracle, MySQL and ClickHouse ... Knowledge of configuration management automation and reducing manual toil * Knowledge of data ...
Business Intelligence Analyst
$140K - $193K/yr
... Kafka/ClickHouse stack. * Validate and reconcile data from disparate systems (MES, tool logs, recipe management, metrology platforms, yield databases) to ensure accurate linkage at the wafer, lot ...
Business Intelligence Analyst
$140K - $193K/yr
... Kafka/ClickHouse stack. * Validate and reconcile data from disparate systems (MES, tool logs, recipe management, metrology platforms, yield databases) to ensure accurate linkage at the wafer, lot ...
Senior Infrastructure Engineer (Core Infra, US)
Palo Alto, CA · On-site +1
$180K - $240K/yr
... ClickHouse, PostgreSQL, or similar technologies. * Experience with observability tools such as Grafana, Prometheus, or VictoriaMetrics. * Understanding of Vault or other secrets management systems.
Quick apply
Senior Infrastructure Engineer (Core Infra, US)
Palo Alto, CA · On-site +1
$180K - $240K/yr
... ClickHouse, PostgreSQL, or similar technologies. * Experience with observability tools such as Grafana, Prometheus, or VictoriaMetrics. * Understanding of Vault or other secrets management systems.
Sr. Database Reliability Engineer
Hawthorne, CA · On-site
$155K - $205K/yr
Ownership of the health, uptime and performance of all PostgreSQL/Oracle/MySQL & Clickhouse ... Knowledge of configuration management automation and limiting the amount of manual toil * Knowledge ...
Sr. Database Reliability Engineer
Hawthorne, CA · On-site
$155K - $205K/yr
Ownership of the health, uptime and performance of all PostgreSQL/Oracle/MySQL & Clickhouse ... Knowledge of configuration management automation and limiting the amount of manual toil * Knowledge ...
Senior Cloud Engineer
San Francisco, CA · On-site
$70.75 - $92/hr
... manage cloud infrastructure across multiple platforms. You should have extensive backend ... Design, develop, and maintain database schemas for Postgres, Clickhouse, and other databases.
Senior Cloud Engineer
San Francisco, CA · On-site
$70.75 - $92/hr
... manage cloud infrastructure across multiple platforms. You should have extensive backend ... Design, develop, and maintain database schemas for Postgres, Clickhouse, and other databases.
Software Engineer, Infrastructure
San Francisco, CA · On-site
$203K - $241K/yr
Infra that users feel - own our core systems (PostgreSQL, ClickHouse, async job processing, LLM ... When do we add Redis, Kafka, a managed vector store? You make the call, ship the migration, and ...
Software Engineer, Infrastructure
San Francisco, CA · On-site
$203K - $241K/yr
Infra that users feel - own our core systems (PostgreSQL, ClickHouse, async job processing, LLM ... When do we add Redis, Kafka, a managed vector store? You make the call, ship the migration, and ...
AI Context & Harness Product Manager
San Francisco, CA · On-site +1
$100K - $250K/yr
Hercules is looking for a senior or staff level product manager to own our AI coding agent. What ... Typescript, React, AI SDK, Postgres, Redis, Clickhouse, Tailwind, IaC (CDK, Terraform), Docker ...
AI Context & Harness Product Manager
San Francisco, CA · On-site +1
$100K - $250K/yr
Hercules is looking for a senior or staff level product manager to own our AI coding agent. What ... Typescript, React, AI SDK, Postgres, Redis, Clickhouse, Tailwind, IaC (CDK, Terraform), Docker ...
Sr. Database Reliability Engineer
Hawthorne, CA · On-site
$155K - $205K/yr
Ownership of the health, uptime and performance of all PostgreSQL/Oracle/MySQL & Clickhouse ... Knowledge of configuration management automation and limiting the amount of manual toil * Knowledge ...
Sr. Database Reliability Engineer
Hawthorne, CA · On-site
$155K - $205K/yr
Ownership of the health, uptime and performance of all PostgreSQL/Oracle/MySQL & Clickhouse ... Knowledge of configuration management automation and limiting the amount of manual toil * Knowledge ...
Devops Engineer
Newport Beach, CA · On-site
$56.50 - $77.50/hr
Manage and monitor cloud environments (Azure preferred; AWS or GCP acceptable) * Collaborate with ... Experience with Microsoft SQL Server and/or ClickHouse * Understanding of networking and security ...
Devops Engineer
Newport Beach, CA · On-site
$56.50 - $77.50/hr
Manage and monitor cloud environments (Azure preferred; AWS or GCP acceptable) * Collaborate with ... Experience with Microsoft SQL Server and/or ClickHouse * Understanding of networking and security ...
Senior Backend Software Engineer, AI Observability & Evals Platform (LangSmith)
San Francisco, CA · On-site
... stack engineers, product managers, and product designers on rolling out new features ... Clickhouse), and cloud platforms (AWS, GCP, Azure) • Strong communication skills, with the ...
Senior Backend Software Engineer, AI Observability & Evals Platform (LangSmith)
San Francisco, CA · On-site
... stack engineers, product managers, and product designers on rolling out new features ... Clickhouse), and cloud platforms (AWS, GCP, Azure) • Strong communication skills, with the ...
Senior Product Manager - Interactive Analytics
Menlo Park, CA · On-site
$200K - $287K/yr
As the Product Manager for Interactive Analytics, you will set the bar for what "fast" means for ... analytics (ClickHouse, Apache Pinot, Materialize, Rockset, SingleStore, or similar); OLTP ...
Senior Product Manager - Interactive Analytics
Menlo Park, CA · On-site
$200K - $287K/yr
As the Product Manager for Interactive Analytics, you will set the bar for what "fast" means for ... analytics (ClickHouse, Apache Pinot, Materialize, Rockset, SingleStore, or similar); OLTP ...
Lead Data Engineer, Analytics & Insights
$160K - $210K/yr
Design and operate the ClickHouse-backed warehouse that powers operational, business, and autonomy ... Experience in robotics, IoT, fleet management, automotive, or any domain where diverse physical ...
Quick apply
Lead Data Engineer, Analytics & Insights
$160K - $210K/yr
Design and operate the ClickHouse-backed warehouse that powers operational, business, and autonomy ... Experience in robotics, IoT, fleet management, automotive, or any domain where diverse physical ...
Devops Engineer
Newport Beach, CA · On-site
$56.75 - $77.75/hr
Manage and monitor cloud environments (Azure preferred; AWS or GCP acceptable) * Collaborate with ... Experience with Microsoft SQL Server and/or ClickHouse * Understanding of networking and security ...
Quick apply
Devops Engineer
Newport Beach, CA · On-site
$56.75 - $77.75/hr
Manage and monitor cloud environments (Azure preferred; AWS or GCP acceptable) * Collaborate with ... Experience with Microsoft SQL Server and/or ClickHouse * Understanding of networking and security ...
Manager Clickhouse information
What is the difference between Manager Clickhouse vs Data Engineer?
| Aspect | Manager Clickhouse | Data Engineer |
|---|---|---|
| Primary Focus | Oversees Clickhouse database management, team coordination, and strategy | Designs, develops, and maintains data pipelines and infrastructure |
| Required Skills | Database administration, team leadership, project management | ETL processes, SQL, programming, data modeling |
| Certifications | Database management, cloud certifications, project management | Data engineering, cloud, and programming certifications |
| Work Environment | Management, strategic planning, team collaboration | Technical development, coding, data architecture |
The Manager Clickhouse primarily focuses on overseeing Clickhouse database operations and leading teams, while a Data Engineer concentrates on building and maintaining data pipelines and infrastructure. Both roles require technical knowledge, but the Manager Clickhouse emphasizes management skills, whereas the Data Engineer emphasizes technical development and data architecture.
Full-time
Re-posted yesterday
Nvidia rating
9.6
Based on 17 frontline employees who took The Breakroom Quiz
8th of 242 rated software companies
Job description
You will work across the NVIDIA AI software stack with teams focused on inference serving, model optimization, distributed systems, GPU performance, and production operations. This is a highly cross-functional role for someone who understands deep learning systems, observability, and large-scale software platforms, and who is excited about building agentic workflows that help teams reason over complex telemetry and performance data.
What You'll Be Doing:
- Lead, mentor, and grow a team building an agentic platform for monitoring and improving large-scale LLMs and VLMs in production.
- Build systems that collect, correlate, and analyze telemetry across inference servers, GPUs, schedulers, model runtimes, and customer-facing APIs.
- Develop agentic workflows that help engineers identify root causes, explain regressions, and recommend performance optimizations.
- Collaborate with internal customers and business units to align priorities and deliver production grade platform capabilities.
What We Need To See:
- BSc, MS, or PhD in Computer Science, Computer Engineering, or equivalent experience.
- 8+ years of relevant software engineering experience, including 3+ years in engineering management or technical leadership.
- Experience leading software engineering teams building large-scale distributed systems, observability platforms, ML infrastructure, or production AI systems.
- Strong understanding of LLM/VLM inference systems, deployment patterns, and production performance challenges.
- Experience with logs, metrics, traces, profiling, alerting, dashboards, or incident/debugging workflows.
- Strong programming, debugging, performance analysis, and test design skills.
- Ability to work across organizations and align technical priorities with product and business goals.
- Excellent communication and collaboration skills.
Ways To Stand Out From The Crowd:
- Background in GPU performance analysis, distributed inference, model serving optimization, or reliability engineering.
- Experience building observability or telemetry platforms for AI, ML, cloud, or distributed infrastructure.
- Experience with OpenTelemetry, Prometheus, Grafana, Jaeger, ClickHouse, Elastic, or similar observability tools.
- Experience building agentic systems that reason over logs, traces, performance data, incidents, or operational workflows.
- Hands-on experience with production GenAI serving systems and metrics such as TTFT, TPOT, throughput, queueing delay, GPU utilization, KV cache pressure, error rates, and cost per token.
#LI-Hybrid
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD for Level 3, and 272,000 USD - 431,250 USD for Level 4.
You will also be eligible for equity and benefits.
Applications for this job will be accepted at least until July 20, 2026.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
About Nvidia
Sourced by ZipRecruiter
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology--and amazing people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what's never been done before takes vision, innovation, and the world's best talent.
Industry
Computer and electronic product manufacturing
Company size
10,000+ Employees
Headquarters location
Santa Clara, CA, US
Year founded
1993