Job Title: Site Reliability Engineer
Location: St. Louis, MO (Hybrid, 3 days onsite/week)
Duration: Long Term Contract
Overview:
Gateway Platforms team is looking for a Site Reliability Engineer to drive our customer experience strategy forward by consistently innovating and problem-solving. The ideal candidate is passionate about the customer experience journey, highly motivated, intellectually curious, analytical, and possesses an entrepreneurial mindset.
Role:
The role of Site Reliability Engineer is to be the production readiness steward for the platform. This is accomplished by closely partnering with developers to design, build, implement, and support technology services. A business operations engineer will ensure operational criteria like system availability, capacity, performance, monitoring, self-healing, and deployment automation are implemented throughout the delivery process. Business Operations plays a key role in leading the DevOps transformation at Client through our tooling and by being an advocate for change and standards throughout the development, quality, release, and product organizations.
We accomplish this transformation through supporting daily operations with a hyper focus on triage and then root cause by understanding the business impact of our products. The goal of every biz ops team is to shift left to be more proactive and upfront in the development process, and to proactively manage production and change activities to maximize customer experience, and increase the overall value of supported applications. Biz Ops teams also focus on risk management by tying all our activities together with an overarching responsibility for compliance and risk mitigation across all our environments. A biz ops focus is also on streamlining and standardizing traditional application specific support activities and centralizing points of interaction for both internal and external partners by communicating effectively with all key stakeholders.
Ultimately, the role of SRE is to align Product and Customer Focused priorities with Operational needs. We regularly review our run state not only from an internal perspective, but also understanding and providing the feedback loop to our development partners on how we can improve the customer experience of our applications.
• Engage in and improve the whole lifecycle of services-from inception and design, through deployment, operations, and refinement.
• Analyse ITSM activities of the platform and provide feedback loop to development teams on operational gaps or resiliency concerns
• Support services before they go live through activities such as system design consulting, capacity planning and launch reviews.
• Maintain services once they are live by measuring and monitoring availability, latency and overall system health.
• Scale systems sustainably through mechanisms like automation and evolve systems by pushing for changes that improve reliability and velocity.
• Support the application CI/CD pipeline for promoting software into higher environments through validation and operational gating, and lead Mastercard in DevOps automation and best practices.
• Practice sustainable incident response and blameless postmortems.
• Take a holistic approach to problem solving, by connecting the dots during a production event thru the various technology stack that makes up the platform, to optimize mean time to recover
• Collaborate with a global team spread across tech hubs in multiple geographies and time zones
• Share knowledge and mentor junior resources.
• Develop and maintain automation pipelines for certificate renewal, traffic routing, alerting, and compliance reporting using tools like Ansible, Venafi.
• Drive improvements in ITSM and DQ SLOs, ensuring timely CRQ status updates and incident closure.
• Lead initiatives for Safety & Soundness and Operational Excellence across quarterly EPICs, covering areas such as PCI compliance, threat/toil management, self-healing, and ITSM defect resolution
All About You
• Background in operational resiliency and self-healing systems.
• Understanding of two factor authentication.
• Strong documentation and communication skills.
• Strong understanding and experience in implementing NGINX configuration.
• Intermediate understanding of Active Directory (Users / Groups), SAML, LTPA, SSO, Oauth.
• Understanding of DEVOPS technologies like Chef, Jenkins, Groovy, shell scripting, bitbucket, GIT.
• Experience in working with or implementing automation workflows and/or scripting development.
• Understanding of:
o Client-server relationships
o Network concepts (Layer 1 to Layer 3)
o Stack trace analysis (TCP dumps, heap dumps, CPU/memory analysis, thread dumps).
o Load balancers and application firewalls.
o Operating System navigation.
o Logging and monitoring methods, standards, and tools.
o High availability and business continuity planning
o Caching concepts
o Configuration management
• Awareness of security implementations, certificate management lifecycle, mutual TLS, SSL handshake, SSH keys, symmetric and asymmetric encryptions.
o Experience with AWS infrastructure and secure access practices.
o Familiarity with ITSM processes, compliance frameworks, and incident management.
o Excellent communication and collaboration skills across cross-functional teams