What We're Looking ForThe Training & Experimentation team at Lightning AI builds the platform that enables developers to experiment at scale and train their own intelligence. Whether customers are training foundation models, fine-tuning open-source models, launching distributed training jobs, or iterating on experiments, this team creates the infrastructure and developer experience that powers every stage of the AI lifecycle.
We're looking for a Senior Software Engineer to help design and scale the backend systems that make AI development faster, more reliable, and easier to use. You'll work across distributed training infrastructure, workload orchestration, experiment management, developer tooling, and platform APIs while partnering closely with our Managed Infrastructure, Core Platform, and Optimized Compute teams.
This is an opportunity to build systems that support real-world AI workloads at scale while shaping the developer experience used by researchers, startups, and enterprise AI teams around the world.
This role is based in one of our San Francisco, NYC, or London office hubs, with a minimum of 2 in-office days per week and occasional team and company offsites. We are not able to provide visa sponsorship for this position at this time.
What You'll Do- Design and build backend services that power AI agent orchestration and execution.
- Develop scalable APIs and platform capabilities for tool use, workflow orchestration, memory, and state management.
- Build reliable infrastructure that enables long-running, distributed AI workflows.
- Partner with product, research, and platform engineering teams to bring new agent capabilities into production.
- Improve platform reliability, observability, and performance as customer workloads continue to scale.
- Evaluate and strengthen technical architecture, engineering processes, and developer tooling.
- Maintain high standards for software quality through testing, automation, continuous delivery, and proactive management of technical debt.
- Mentor engineers on system design, distributed systems, and backend engineering best practices
What You'll Need- Extensive experience building scalable backend systems using Go, Python, Rust, or similar languages.
- Strong understanding of distributed systems, asynchronous processing, APIs, and cloud-native application architecture.
- Experience designing and operating production services at scale.
- Hands-on experience building backend platforms, developer tools, or infrastructure products.
- Familiarity with cloud platforms such as AWS, GCP, or Azure and container technologies including Docker and Kubernetes.
- Proven ability to take ownership of large technical initiatives from architecture through production.
- Commitment to software quality, automated testing, observability, and continuous delivery.
- Strong communication skills and the ability to collaborate effectively across engineering, product, research, and design teams.
- Ability to thrive in fast-moving, ambiguous environments while balancing technical excellence with product impact.
Compensation