At American Express, we are transforming how software is built, deployed, and operated at enterprise scale. The Site Reliability Engineering (SRE) organization is at the center of this transformation, building the engineering capabilities that enable secure software delivery, resilient platforms, intelligent operations, and exceptional developer experiences.
As a Staff Engineer, you will be a senior technical leader responsible for advancing the reliability, scalability, and operational excellence of critical technology platforms. You will work across engineering organizations to solve complex production challenges through software engineering, platform innovation, automation, and AI-powered operational capabilities.
This is a highly influential individual contributor role where success is measured by the systems you build, the engineering practices you shape, and the enterprise impact you create.
Role Summary
As a Staff Engineer within the Site Reliability Engineering organization, you will lead the evolution of enterprise reliability engineering by building scalable platform capabilities that improve system resilience, engineering productivity, and operational efficiency.
You will partner closely with Platform Engineering, Infrastructure Engineering, Security Engineering, Application Development, and Enterprise Architecture teams to modernize how software is delivered, observed, secured, and operated across the enterprise.
This role requires deep expertise in distributed systems, cloud infrastructure, software engineering, DevSecOps, observability, automation, and AI-enabled operations. You will drive engineering solutions that reduce operational complexity, improve service reliability, eliminate manual toil, and accelerate the organization's journey toward autonomous operations.
At American Express, our culture is built on a 175-year history of innovation, shared values and Leadership Behaviors, and an unwavering commitment to back our customers, communities, and colleagues. From delivering differentiated products to providing world-class customer service, we operate with a strong risk mindset, ensuring we continue to uphold our brand promise of trust, security, and service.
As part of Team Amex, you'll experience our powerful backing with comprehensive support for your holistic well-being and many opportunities to learn new skills, develop as a leader, and grow your career. Here, your voice and ideas matter, your work makes an impact, and together, you will help us define the future of American Express.
We back you with benefits that support your holistic well-being so you can be and deliver your best. This means caring for you and your loved ones' physical, financial, and mental health, as well as providing the flexibility you need to thrive personally and professionally:
- Competitive base salaries
- Bonus incentives
- 6% Company Match on retirement savings plan
- Free financial coaching and financial well-being support
- Comprehensive medical, dental, vision, life insurance, and disability benefits
- Flexible working model with hybrid, onsite or virtual arrangements depending on role and business need
- 20+ weeks paid parental leave for all parents, regardless of gender, offered for pregnancy, adoption or surrogacy
- Free access to global on-site wellness centers staffed with nurses and doctors (depending on location)
- Free and confidential counseling support through our Healthy Minds program
- Career development and training opportunities
For a full list of Team Amex benefits, visit our Colleague Benefits Site.
American Express is an equal opportunity employer and makes employment decisions without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran status, disability status, age, or any other status protected by law. American Express will consider for employment all qualified applicants, including those with arrest or conviction records, in accordance with the requirements of applicable state and local laws, including the California Fair Chance Act, the Los Angeles County Fair Chance Ordinance for Employers, and the City of Los Angeles' Fair Chance Initiative for Hiring Ordinance. For positions covered by federal and/or state banking regulations, American Express will comply with such regulations as it relates to the consideration of applicants with criminal convictions.
We back our colleagues with the support they need to thrive, professionally and personally. That's why we have Amex Flex, our enterprise working model that provides greater flexibility to colleagues while ensuring we preserve the important aspects of our unique in-person culture. Depending on role and business needs, colleagues will either work onsite, in a hybrid model (combination of in-office and virtual days) or fully virtually.
US Job Seekers - Click to view the "Know Your Rights" poster. If the link does not work, you may access the poster by copying and pasting the following URL in a new browser window: https://www.eeoc.gov/poster.
The below represents the expected salary range for this job requisition. Ultimately, in determining your pay, we'll consider your location, experience, and other job-related factors.
Minimum Qualifications- Bachelor's degree in Computer Science, Engineering, or related technical discipline, or equivalent practical experience.
- 12+ years of experience in software engineering, Site Reliability Engineering, cloud infrastructure, platform engineering, or distributed systems.
- Demonstrated experience designing and operating highly available, large-scale distributed systems.
- Strong software engineering background with proficiency in one or more modern programming languages such as Java, Go, Python, or TypeScript.
- Deep understanding of cloud-native architectures, Kubernetes, Infrastructure as Code, CI/CD, observability, and engineering automation.
- Experience applying SRE principles, including SLOs, error budgets, incident management, production readiness, and operational excellence.
- Proven ability to influence technical direction across multiple engineering organizations.
Preferred QualificationsExperience with several of the following technologies and practices:
Cloud & Infrastructure
- AWS
- Kubernetes/OpenShift
- Docker
- Terraform
- CloudFormation
- Ansible
Platform Engineering
- GitHub Actions
- Jenkins
- GitLab CI
- Argo CD
Observability & Reliability
- OpenTelemetry
- Prometheus
- Grafana
- Splunk
- Dynatrace
- Elastic
- Chaos Engineering
- Capacity Engineering
- Performance Engineering
Software Engineering
- Java
- Go
- Python
- Event-driven architectures
- Kafka
- Distributed systems design
AI & Automation
- Generative AI
- Agentic AI
- AIOps
- Workflow orchestration
- Intelligent automation platforms
- AI-assisted engineering tools
Leadership ExpectationsSuccessful Staff Engineers at American Express:
- Solve complex engineering problems through software, automation, and scalable platform solutions.
- Balance innovation with operational excellence, resilience, and security.
- Influence engineering direction through technical expertise rather than organizational authority.
- Remain hands-on by designing systems, reviewing critical implementations, and developing prototypes.
- Build consensus across engineering organizations through collaboration and technical credibility.
- Mentor engineers and cultivate a culture of learning, ownership, and engineering excellence.
- Continuously evaluate emerging technologies that improve reliability, developer productivity, and operational efficiency.
Why Join Us?As a Staff Engineer in the Site Reliability Engineering organization, you will help shape the future of engineering at American Express. Your work will influence how thousands of engineers build, deploy, secure, observe, and operate software across one of the world's largest financial technology platforms.
You will have the opportunity to build foundational engineering capabilities, modernize enterprise software delivery, advance AI-powered operations, and drive the next generation of autonomous engineering-all while improving the reliability and resilience of services that millions of customers depend on every day.
Employment eligibility to work with American Express in the United States is required as the company will not pursue visa sponsorship for these positions.
Key ResponsibilitiesLead Enterprise Reliability Engineering
Drive the technical direction for reliability engineering across critical enterprise platforms.
You will:
- Define and evolve enterprise-wide Site Reliability Engineering practices and standards.
- Champion Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets to improve service reliability and engineering accountability.
- Establish engineering patterns for high availability, scalability, resiliency, and disaster recovery.
- Lead architecture and design reviews for mission-critical platforms and distributed systems.
- Partner with engineering teams to improve production readiness, operational maturity, and service resilience.
- Drive adoption of proactive reliability engineering practices, including capacity planning, performance optimization, failure testing, and resilience validation.
- Build Software That Improves Operations
Apply software engineering to eliminate operational complexity and improve engineering productivity.
You will:
- Design and develop reusable platforms, frameworks, and automation that reduce operational toil.
- Build engineering capabilities that simplify production operations and enable self-service experiences.
- Develop reference implementations and engineering libraries adopted across multiple organizations.
- Improve deployment safety through automated validation, policy enforcement, and release controls.
Leverage modern programming languages and cloud-native technologies to solve complex operational problems at scale.
Modernize Software Delivery & Engineering PlatformsEnable secure, reliable, and efficient software delivery across the enterprise.
You will:
- Drive maturity of enterprise CI/CD platforms through standardized pipelines and reusable engineering capabilities.
- Embed security, reliability, testing, and compliance directly into software delivery workflows.
- Advance Platform Engineering and Internal Developer Platforms (IDPs) that improve developer experience and engineering velocity.
- Improve deployment reliability using progressive delivery, automated verification, rollback strategies, and policy-as-code.
Partner with application teams to simplify software delivery while improving production stability.
Advance Intelligent Operations & Autonomous EngineeringLead the evolution of modern operations through AI-driven engineering and intelligent automation.
You will:
- Build next-generation operational capabilities using AI, machine learning, and intelligent automation.
- Integrate Generative AI and agentic AI into incident management, diagnostics, troubleshooting, and engineering workflows.
- Design self-healing operational capabilities that automate detection, analysis, and remediation of production issues.
- Improve operational decision-making through predictive analytics, anomaly detection, and intelligent event correlation.
Advance the organization's journey toward autonomous operations while maintaining governance and operational safety.
Improve Production Security & Operational ResiliencePartner with Security Engineering to improve the resilience of production systems through secure engineering practices.
Integrate automated vulnerability detection and remediation into software delivery and operational workflows.
Build scalable engineering solutions that reduce vulnerability remediation time and improve enterprise security posture.
Advance secure-by-default engineering practices across infrastructure and application platforms.
Improve software supply chain security through automated controls and continuous compliance.
Reduce operational risk through engineering automation and policy-driven governance.
Drive Enterprise AutomationReduce operational toil through intelligent automation and reusable engineering solutions.
Identify repetitive operational activities and replace them with scalable automation.
Build reusable automation frameworks supporting infrastructure, application operations, compliance, and engineering workflows.
Standardize operational runbooks through workflow orchestration and policy-driven automation.
Improve engineering efficiency by enabling self-service operational capabilities.
Measure automation effectiveness through reductions in manual effort, operational risk, and incident resolution time.
Provide Technical LeadershipAs a Staff Engineer, your influence extends beyond any single team.
Lead enterprise-wide engineering initiatives spanning multiple organizations.
Influence architectural decisions through technical expertise and engineering excellence.
Mentor senior engineers and technical leaders across the organization.
Prototype emerging technologies and evaluate their applicability to enterprise engineering challenges.
Drive technical alignment through collaboration, technical design reviews, and engineering standards.
Foster a culture of continuous improvement, operational excellence, and engineering innovation.
What Success Looks LikeSuccess in this role will be measured through tangible engineering outcomes, including:
Improved service availability, reliability, and platform resilience.
Increased adoption of SLO-driven engineering and operational maturity practices.
Reduced Mean Time to Detect (MTTD) and M...