Contract Details
- Work Mode: 100% Remote (US-based)
- Location: Herndon, VA
- Schedule: 40 hours/week
- Duration: 08/17/2026 08/16/2027
- Type: Contract with potential to convert to full-time after ~12 months (not guaranteed)
About the Opportunity
Seeking a Site Reliability Engineer to ensure availability, performance, scalability, and security for mission-critical, cloud-hosted search and analytics services built on OpenSearch. You will focus on reliability engineering, operations, automation, and continuous improvement for distributed platforms, working within a diverse, globally distributed team.
Key Responsibilities
- Provision, build, deploy, monitor, operate, and support cloud services in a global team environment.
- Architect, build, deploy, and maintain high performance OpenSearch clusters and platforms from the ground up.
- Optimize OpenSearch for high availability, resiliency, scalability, security, and performance.
- Monitor and troubleshoot cluster health, node performance, indexing throughput, search latency, shard allocation/replication, and storage utilization.
- Analyze and resolve operational issues across infrastructure, platform, and application layers; lead incident response, RCA, and remediation.
- Maintain integrity and security of servers, systems, and OpenSearch platform infrastructure.
- Support lifecycle activities: installation, configuration, upgrades/patching, backup/restore, and disaster recovery.
- Develop and maintain monitoring policies, alerting standards, runbooks, and support procedures.
- Automate testing, deployment, scaling, recovery, and operational workflows for OpenSearch and related cloud services.
- Plan capacity for compute, memory, storage, and network; partner with engineering to enhance reliability and operational readiness.
- Support log ingestion, index management, lifecycle/retention, and search performance tuning.
- Participate in an on-call rotation; support occasional weekend/after-hours needs.
Required Qualifications
- US citizenship required; dual citizenship not permitted.
- 8 years of experience in SRE/DevOps/cloud operations with distributed systems.
- Proven, hands-on experience designing, building, deploying, operating, and optimizing OpenSearch clusters from scratch in production.
- Expert-level Kubernetes experience (operations, troubleshooting, management, configuration of complex services).
- Deep OpenSearch administration: cluster architecture, performance tuning, scaling, upgrades, and troubleshooting; index/shard/replica strategy; sizing; snapshot/restore; backup/DR.
- Strong Linux expertise (SUSE and Ubuntu).
- Expertise with Git and Concourse (pipeline setup, management, troubleshooting).
- Experience with Kafka and Zookeeper; strong automation for testing, deployment, scalability, and cloud service management.
- Experience building/implementing/supporting cloud monitoring and observability; solid knowledge of cloud computing, infrastructure operations, databases, web services, networking, virtualization, and internet protocols.
- Security fundamentals for SaaS multi-tenant application systems; excellent communication and prioritization skills; ability to multitask.
Preferred Qualifications
- AWS experience (e.g., Route 53, EC2, S3, CloudWatch, DynamoDB, RDS, IAM, ACM, KMS, VPC); experience deploying/operating OpenSearch in AWS.
- Experience with Cloud Foundry environments.
- Experience with Jenkins, Chef, and/or Terraform.
- Experience with Prometheus and Grafana.
- Background with log ingestion pipelines, index lifecycle management, retention strategies, and search platform security controls.
- Familiarity with capacity forecasting, performance benchmarking, and resilience testing for distributed search platforms.
Work Environment
- Collaborative, globally distributed team with cross-training opportunities.
- Participation in an on-call rotation and occasional after-hours/weekend support.
#ZR