Job Summary:
VITURE is the #1 XR glasses brand in the US, aiming to create the first great AI interface you wear. The Site Reliability Engineer will design and maintain scalable cloud infrastructure for intelligent eyewear, ensuring high reliability and performance for users worldwide.
Responsibilities:
• Design, deploy, and manage scalable infrastructure across mainstream cloud platforms to support our high-traffic AI services (e.g., LLM inference pipelines, real-time voice, and spatial computing backends).
• Establish and execute incident management protocols, participate in on-call rotations, and lead blameless post-mortems to continuously reduce Mean Time to Recovery (MTTR) and improve system reliability.
• Champion Infrastructure as Code (IaC) principles. Leverage automation tools (e.g., Terraform, Ansible) to automate provisioning, configuration, and deployments, actively eliminating manual operational toil.
• Build and maintain comprehensive observability platforms (monitoring, logging, tracing) to track real-time resource utilization, define critical metrics, and ensure strict Service Level Objectives (SLOs) are met.
• Collaborate closely with software and AI teams to embed reliability early in the development lifecycle, streamline CI/CD pipelines, and optimize performance across servers and containerized environments.
• Enforce robust security postures by implementing network policies, configuring firewalls, and establishing rigorous backup and Disaster Recovery (DR) strategies.
• Mentor junior engineers and contribute to fostering a culture of engineering excellence and reliability as needed.
Qualifications:
Required:
• Bachelor's degree or above in Computer Science, Electrical Engineering, or a related field; or equivalent practical experience with a demonstrated track record of exceptional work.
• 2+ years of professional experience in an SRE, DevOps, or Cloud Infrastructure role, preferably supporting high-concurrency consumer applications or AI services.
• Strong hands-on experience managing and scaling distributed systems on mainstream public or private cloud ecosystems.
• Solid expertise in containerization (e.g., Docker) and orchestration platforms (e.g., Kubernetes) for deploying complex microservices architectures.
• Proficiency in scripting or programming languages (such as Python, Go, or Shell) to develop custom automation and tooling.
• Experience with modern observability, telemetry, and centralized logging stacks (e.g., Prometheus, Grafana, ELK Stack).
• Excellent communication skills with the ability to articulate technical decisions and collaborate effectively with cross-functional teams.
Preferred:
• Previous experience managing infrastructure for AI/ML workloads (e.g., GPU cluster management, multimodal AI inference deployments).
• Personal project experience is a strong plus — we love seeing what you build for fun. Share your side projects, open-source infrastructure contributions, or passion builds with us.
• Product-minded with a refined standard for reliability and craft — you care deeply about the user experience when systems degrade gracefully, not just whether the servers are running.
Company:
VITURE is a 2C technology company dedicated to developing AR technology, creating trendy XR glasses. Founded in 2021, the company is headquartered in San Francisco, USA, with a team of 201-500 employees. The company is currently Growth Stage.