- System Resilience: Design, build, and maintain highly available, scalable, and secure infrastructure to support our AI-native cybersecurity platform.
- Automation & Tooling: Develop internal tooling and automation to streamline deployment processes, incident response, and capacity planning.
- Performance Engineering: Monitor system performance and proactively identify bottlenecks, optimizing infrastructure for low-latency, high-throughput AI workloads.
- Incident Management: Lead incident response efforts, conduct post-mortems, and implement long-term solutions to prevent recurring reliability issues.
- Infrastructure as Code (IaC): Manage infrastructure via code, driving consistency, auditability, and scalability across our cloud environments (e.g., AWS, GCP).
- Cross-Functional Collaboration: Partner with sibling Engineering teams, Product, and Security teams to ensure reliability is baked into our development lifecycle from concept to production.
Requirements
- Core Engineering: 10+ years of experience in SRE, DevOps, or Software/Systems Engineering, particularly in managing production systems at scale.
- Cloud Infrastructure: Deep expertise in public cloud environments (AWS, GCP, or Azure) and managing services such as Kubernetes (EKS/GKE), networking, and storage.
- Infrastructure as Code: Extensive experience with tools like Terraform, Pulumi, or similar technologies to manage complex infrastructure deployments.
- Observability: Hands-on experience with monitoring, logging, and tracing stacks (e.g., Prometheus, Grafana, ELK, Datadog) to drive data-informed reliability decisions.
- Distributed Systems: Solid understanding of microservices architecture, distributed databases, and event-driven systems.
- Communication: Clear, concise communication skills and a bias for collaborative problem-solving.
- Leadership Alignment: Proven track record of guiding multi-stakeholder initiatives and influencing engineering practices across teams.
- Analytical Rigor: Strong problem-solving, debugging, and analytical skills, especially in high-pressure environments.
- Domain Background: Prior work in cybersecurity, specifically regarding SIEM, EDR, or SOAR infrastructure is nice-to-have.
- AI/ML Infrastructure: Experience supporting infrastructure for large-scale AI/ML workloads (e.g., GPU scheduling, LLM serving optimization) is nice-to-have.
- Startup Mentality: Background driving high-impact engineering initiatives in high-growth startups or enterprise SaaS is nice-to-have.
- Strong familiarity with Agentic Workflows such as Agno, Temporal, etc. is nice-to-have.
Core Competencies
Demonstrates expertise in designing and maintaining scalable, secure cloud infrastructure, with a strong focus on automation, performance engineering, and incident management. Proven ability to collaborate across teams and drive reliability in AI-native cybersecurity platforms.
Highest-signal resume keywords
- 10+ Years Experience in SRE, DevOps, or Software Engineering
- Deep Expertise in AWS, GCP, or Azure
- Extensive Experience with Terraform or Pulumi
- Hands-On Experience with Prometheus, Grafana, or ELK
- Strong Problem-Solving and Analytical Skills
ATS Optimization Keywords Hard Skills
- Infrastructure as Code
- Performance Engineering
- Incident Management
- Cloud Infrastructure Management
- Distributed Systems Understanding
- AI/ML Infrastructure Support
- Automation Development
- Monitoring and Observability
- Capacity Planning
- Microservices Architecture
Soft Skills
- Clear Communication Skills
- Collaborative Problem-Solving
- Leadership in Multi-Stakeholder Initiatives
Industry Keywords
- Cybersecurity
- AI-Native Platforms
- High-Impact Engineering
- Startup Mentality
- Enterprise SaaS
Tools & Technologies
- Kubernetes (EKS/GKE)
- Prometheus
- Grafana
- ELK
- Datadog
- Terraform
- Pulumi
- Agentic Workflows
- SIEM
- EDR
#J-18808-Ljbffr