Job Summary:
Charlie Health is a rapidly growing organization focused on providing personalized, virtual behavioral health treatment. They are seeking a Senior Machine Learning Platform Engineer to define the technical direction and build foundational systems for their ML and AI capabilities, ensuring scalable and reliable delivery of AI-powered features.
Responsibilities:
• Define technical direction for ML/AI infrastructure and make build-vs-buy decisions as the founding platform engineer
• Design and operate multi-vendor AI infrastructure supporting client-facing and clinician-facing LLM applications across multiple LLM providers
• Design, build, and operate production model serving systems; maintain infrastructure as code for reproducible ML environments, training pipelines, and deployment workflows
• Develop high-performance GPU inference pipelines with low latency and high availability
• Own the multimodal data pipeline layer—manage ingestion, processing, and serving of text, audio, and structured clinical data for ML and AI systems
• Create reliable infrastructure for agentic AI systems, including orchestration, monitoring, evaluation, and observability tooling
• Build developer tooling that accelerates data science and ML engineering workflows across the organization
• Own AI observability—build monitoring, alerting, and debugging capabilities for production ML systems
• Partner with ML engineers, data scientists, and product teams to understand infrastructure needs and translate them into scalable platform capabilities
• Foster a culture of collaboration and learning across engineering, product, and design through mentoring, documentation, presentations, and knowledge sharing
• Participate in our on-call rotation to ensure model serving uptime, pipeline reliability, and infrastructure health
Qualifications:
Required:
• 4+ years of professional experience in software engineering, with at least 2 years focused on ML infrastructure, ML platform, or AI systems engineering
• Strong software engineering fundamentals in Python and deep infrastructure expertise
• Familiarity with cloud ML services (AWS SageMaker, GCP Vertex AI, or similar) and CI/CD for ML pipelines
• Experience with infrastructure as code (Terraform, Pulumi, or similar) and container orchestration (Kubernetes, ECS)
• Excellent at managing ambiguity—able to break down big, messy problems into smaller parts with tractable solutions and clear iterations
• Growth mindset and sense of humor; you welcome feedback, adapt quickly in a fast-paced environment, and foster a culture of learning and fun
Preferred:
• Experience building evaluation and observability systems for LLM-based or agentic AI applications is a plus
• Experience with a systems language (Go, Rust, or C++) is a plus
Company:
Virtual behavioral health clinic for high acuity youth Founded in 2020, the company is headquartered in Bozeman, USA, with a team of 1001-5000 employees. The company is currently Late Stage.