Job Summary:
SpaceX is a pioneering aerospace manufacturer and space transport services company focused on enabling human life on Mars. They are seeking a Site Reliability Engineer to build and maintain mission-critical platforms that accelerate vehicle software delivery and operations. The role involves deploying and maintaining products, collaborating with software engineers, and providing end-user support.
Responsibilities:
• Deploy, upgrade, operate, maintain, and scale our suite of mission-critical products and services
• Manage our underlying infrastructure as code and use modern observability tools to provide a complete picture of application health
• Closely collaborate with software engineers to design and build highly operable, maintainable, and testable systems
• Engage in and improve the entire software development lifecycle — from inception and design through deployment, operation, and continuous refinement
• Practice sustainable incident response and blameless postmortems
• Provide high-quality end-user support to vehicle software engineers
• Participate in the team’s on-call rotation
• Identify and eliminate performance bottlenecks using measurement and creative engineering
Qualifications:
Required:
• Bachelor’s degree in computer science, information systems, or an engineering discipline; OR 3+ years of professional experience in SRE or DevOps in lieu of a degree
• 1+ years of experience with Python and Python-based development frameworks
• Experience with Linux operating systems
• Must be able to work extended hours and weekends as needed
Preferred:
• Experience with build systems (Bazel, Buck, Make, etc.)
• Experience with both container and virtualization technologies (Docker, Kubernetes, vSphere, QEMU, KVM, etc.)
• Experience with databases and data modeling (Postgres, MySQL, ClickHouse, etc.)
• Experience with infrastructure as code (IaC) tools for managing fleets of servers
• Experience with Terraform, Ansible, Puppet, or similar automation frameworks
• Knowledge of the technologies that predate and underpin modern cloud infrastructure, with the ability to translate high-level developer experiences into specific implementations from first principles
• Ability to work with mission-critical and sensitive systems with appropriate urgency and care
• Ability to communicate effectively with customers, peers, and management in both formal and informal settings
• Experience with full-stack development (the team primarily uses Python, JavaScript, and C#; end users primarily use C++)
Company:
SpaceX develops and operates rockets, satellite networks, and AI infrastructure including launch, connectivity, and cloud services. Founded in 2002, the company is headquartered in Hawthorne, USA, with a team of 1001-5000 employees. The company is currently Late Stage.