Job Title: CockroachDB Senior Engineer
Location: Sunnyvale, CA (Onsite)
Seniority on the skill/s required on this requirement: Senior
Estimated Duration: Full Time
Work authorization: USC, GC, GC EAD Only
Interview Process: Video Interviews
Job Summary:
We are seeking a Senior CockroachDB Engineer to design, deploy, operate, and scale multi-region CockroachDB clusters in production environments. This role is critical to ensuring high availability, fault tolerance, and data consistency across globally distributed database clusters, while driving performance optimization, disaster recovery planning, and operational excellence.
Responsibilities:
- Design, deploy, operate, and scale multi-region CockroachDB clusters in production environments
- Ensure high availability, fault tolerance, and data consistency for globally distributed clusters
- Monitor cluster health, latency, replication status, and resource utilization using observability tools
- Perform capacity planning and proactive scaling for future growth
- Troubleshoot complex database and infrastructure issues, including:
- Node failures
- Network partitions
- Leaseholder and range imbalance
- Replication lag
- Hotspotting
- High latency / throughput bottlenecks
- Design disaster recovery strategies (multi-region, backup/restore, failover/fallback)
- Implement and test backup, restore, and point-in-time recovery processes
- Automate provisioning, scaling, patching, and upgrades of CRDB clusters
- Perform rolling upgrades with zero or near-zero downtime
- Optimize SQL query performance and database schema efficiency
- Create operational runbooks, SOPs, and on-call playbooks for CRDB
- Participate in on-call rotations and incident response for production clusters
Requirements:
- Proven senior-level experience designing, deploying, and operating CockroachDB (CRDB) clusters in production
- Strong background in distributed database systems, high availability, and fault-tolerant architecture
- Hands-on experience with disaster recovery planning, backup/restore, and point-in-time recovery
- Experience with observability/monitoring tools for tracking cluster health, latency, and replication
- Strong troubleshooting skills across database, network, and infrastructure layers
- Experience automating provisioning, patching, and upgrade processes for database clusters
- Strong SQL query optimization and schema design skills
- Must be authorized to work in the US (USC, GC, or GC EAD only)
- Willingness to participate in on-call rotations
Mindtris, a minority women-owned enterprise, is at the forefront of digital transformation, technology excellence, and business growth solutions. Specializing in talent mobilization and innovation, we are dedicated to enhancing customer experiences across diverse sectors such as Information Technology, Telecommunications, Healthcare, Engineering, and the Public sector. With a focus on deploying top-tier talent and fostering innovation, we empower businesses to thrive and excel in a rapidly evolving digital landscape, helping them reach new heights of success.
Mindtris is committed to fostering workforce diversity and is proud to be an equal opportunity employer.