Job Summary:
Obsidian Security is the leading SaaS security platform, trusted by global enterprises like Snowflake, T-Mobile, and Algolia. They are seeking a Site Reliability Engineer to support and maintain the service quality of their SaaS security platform while addressing challenges around scalability and reliability.
Responsibilities:
โข Support and maintain the service quality of our customer-facing SaaS security platform
โข Address complex challenges around scalability, reliability, observability, and cost efficiency
โข Collaborate with Engineering teams to maintain and enhance Helm charts, application deployment, monitoring and CI/CD pipelines
โข Embed into the engineering team so that you understand the application deeply.
โข Define service verification strategies and implement them as part of the CI/CD process to meet SLAs
โข Improve developer experience by optimizing CI/CD workflows and performance
โข Participate in the on-call rotation, providing 24/7 support in coordination with our global SRE team
โข Monitor, debug, and optimize production infrastructure and services on AWS/GCP
Qualifications:
Required:
โข 3+ years of experience in a DevOps or SRE role supporting SaaS services on GCP and/or AWS
โข Bachelor's degree in Computer Science or related field
โข Strong proficiency in Kubernetes, microservices architecture, Helm, GitLab CI/CD, and ArgoCD, Prometheus, Grafana.
โข Deep understanding of autoscaling, version upgrades, and cloud service optimization
Preferred:
โข Programming experience in at least one language; Golang or Python preferred
โข Bonus if you're familiar with technologies like Kafka, Elasticsearch, PostgreSQL, ScyllaDB, Databricks, Dagster, Sentry, Kong
Company:
Obsidian Security provides SaaS and AI security software that detects threats, manages risks, and secures enterprise applications. Founded in 2017, the company is headquartered in Newport Beach, USA, with a team of 51-200 employees. The company is currently Growth Stage.