ABOUT YOU We are looking for a Site Reliability Engineer (Monetization) who is pragmatic, product-minded, and equally comfortable writing code and running production systems to join our ...
ABOUT YOU We are looking for a Site Reliability Engineer (Monetization) who is pragmatic, product-minded, and equally comfortable writing code and running production systems to join our ...
Site Reliability Engineer
Montreal, QC · On-site +1
About the Role We are looking for a Site Reliability Engineer to join our Network and Security Operations Center (NOC), a team at the heart of platform reliability for mission-critical SaaS ...
Site Reliability Engineer
Montreal, QC · On-site +1
About the Role We are looking for a Site Reliability Engineer to join our Network and Security Operations Center (NOC), a team at the heart of platform reliability for mission-critical SaaS ...
... (SRE) est une discipline orientee production, axee sur l'amelioration de la disponibilite des services systemes, de l'observabilite, de l'evolutivite, de la performance et de la fiabilite des ...
... (SRE) est une discipline orientee production, axee sur l'amelioration de la disponibilite des services systemes, de l'observabilite, de l'evolutivite, de la performance et de la fiabilite des ...
Site Reliability Engineer
Montreal, QC · On-site
About the Role We are looking for a Site Reliability Engineer to join our Network and Security Operations Center (NOC), a team at the heart of platform reliability for mission-critical SaaS ...
Quick apply
Site Reliability Engineer
Montreal, QC · On-site
About the Role We are looking for a Site Reliability Engineer to join our Network and Security Operations Center (NOC), a team at the heart of platform reliability for mission-critical SaaS ...
... (SRE) est une discipline orientee production, axee sur l'amelioration de la disponibilite des services systemes, de l'observabilite, de l'evolutivite, de la performance et de la fiabilite des ...
... (SRE) est une discipline orientee production, axee sur l'amelioration de la disponibilite des services systemes, de l'observabilite, de l'evolutivite, de la performance et de la fiabilite des ...
Senior Site Reliability Engineer
Montreal, QC · Remote
$95K - $100K/yr
We are looking for an experienced Senior Site Reliability Engineer for our client. This is a permanent position that is remote to start with later relocation to Calgary or Winnipeg . Our client is a ...
Quick apply
Senior Site Reliability Engineer
Montreal, QC · Remote
$95K - $100K/yr
We are looking for an experienced Senior Site Reliability Engineer for our client. This is a permanent position that is remote to start with later relocation to Calgary or Winnipeg . Our client is a ...
Senior Site Reliability Engineer
Montreal, QC · Remote
$95K - $100K/yr
We are looking for an experienced Senior Site Reliability Engineer for our client. This is a permanent position that is remote to start with later relocation to Calgary or Winnipeg . Our client is a ...
Quick apply
Senior Site Reliability Engineer
Montreal, QC · Remote
$95K - $100K/yr
We are looking for an experienced Senior Site Reliability Engineer for our client. This is a permanent position that is remote to start with later relocation to Calgary or Winnipeg . Our client is a ...
SRE specialist
Montreal, QC · Hybrid
The SRE will function as part of a special investigations unit that empowers and enables Applicative Support, Infrastructure Support, and the Incident Management team-coaching, guiding, and leading ...
SRE specialist
Montreal, QC · Hybrid
The SRE will function as part of a special investigations unit that empowers and enables Applicative Support, Infrastructure Support, and the Incident Management team-coaching, guiding, and leading ...
SRE specialist
Montreal, QC · Hybrid
The SRE will function as part of a special investigations unit that empowers and enables Applicative Support, Infrastructure Support, and the Incident Management team-coaching, guiding, and leading ...
SRE specialist
Montreal, QC · Hybrid
The SRE will function as part of a special investigations unit that empowers and enables Applicative Support, Infrastructure Support, and the Incident Management team-coaching, guiding, and leading ...
About the Role The Director of Infrastructure & SRE owns the function end-to-end: reliability, security, scalability, and operational governance of TailorCare's infrastructure, plus the team that ...
About the Role The Director of Infrastructure & SRE owns the function end-to-end: reliability, security, scalability, and operational governance of TailorCare's infrastructure, plus the team that ...
Site Reliability Engineer
Montreal, QC · On-site
Nous sommes a la recherche d'un(e) ingenieur(e) en fiabilite des sites (Site Reliability Engineer - SRE) competent(e) pour concevoir, exploiter et ameliorer des systemes de production hautement ...
Site Reliability Engineer
Montreal, QC · On-site
Nous sommes a la recherche d'un(e) ingenieur(e) en fiabilite des sites (Site Reliability Engineer - SRE) competent(e) pour concevoir, exploiter et ameliorer des systemes de production hautement ...
We are looking for an experienced Site Reliability Engineer or Platform Operations Engineer for our client. This is a permanent position that is remote to start with later relocation to Calgary or ...
Quick apply
We are looking for an experienced Site Reliability Engineer or Platform Operations Engineer for our client. This is a permanent position that is remote to start with later relocation to Calgary or ...
We are looking for an experienced Site Reliability Engineer or Platform Operations Engineer for our client. This is a permanent position that is remote to start with later relocation to Calgary or ...
Quick apply
We are looking for an experienced Site Reliability Engineer or Platform Operations Engineer for our client. This is a permanent position that is remote to start with later relocation to Calgary or ...
SRE Associate (Hybrid)
Montreal, QC · Hybrid
We're seeking someone to join our team as a Site Reliability Engineering Specialist to support the reliability, performance, and operational stability of business-critical applications, while helping ...
SRE Associate (Hybrid)
Montreal, QC · Hybrid
We're seeking someone to join our team as a Site Reliability Engineering Specialist to support the reliability, performance, and operational stability of business-critical applications, while helping ...
We are growing SRE capabilities within our Reliability & Production Engineering (RPE) organization as part of the transformation of Morgan Stanley's Technology. In the Technology division, we ...
We are growing SRE capabilities within our Reliability & Production Engineering (RPE) organization as part of the transformation of Morgan Stanley's Technology. In the Technology division, we ...
Sur certains mandats, leadership technique - Piloter le volet fiabilite et infrastructure, guider les equipes clients sur les pratiques SRE et contribuer aux decisions architecturales. * Soutien a ...
Sur certains mandats, leadership technique - Piloter le volet fiabilite et infrastructure, guider les equipes clients sur les pratiques SRE et contribuer aux decisions architecturales. * Soutien a ...
DevOps / SRE Engineer (Remote)
Montreal, QC · On-site +1
To do that we are eager to add a highly skilled DevOps / SRE Engineer Engineer to our incredible team. This is a senior role working alongside the current backend and frontend software engineering ...
Quick apply
DevOps / SRE Engineer (Remote)
Montreal, QC · On-site +1
To do that we are eager to add a highly skilled DevOps / SRE Engineer Engineer to our incredible team. This is a senior role working alongside the current backend and frontend software engineering ...
As a Site Reliability Specialist at Ubisoft Montréal, you will join the IT Games and Studios team ... You'll collaborate with developers, cloud specialists, and infrastructure teams to build resilient ...
Quick apply
As a Site Reliability Specialist at Ubisoft Montréal, you will join the IT Games and Studios team ... You'll collaborate with developers, cloud specialists, and infrastructure teams to build resilient ...
Intermediate DevOPS/SRE
Montreal, QC · On-site
The COO/GTE/EPL/SRE team has members in Paris, Bangalore, and Montreal and is responsible for the production, security, performance, and scalability of all capabilities provided by EPL. WHAT WILL BE ...
Intermediate DevOPS/SRE
Montreal, QC · On-site
The COO/GTE/EPL/SRE team has members in Paris, Bangalore, and Montreal and is responsible for the production, security, performance, and scalability of all capabilities provided by EPL. WHAT WILL BE ...
Site Reliability Specialist
Sorel-tracy, QC · Hybrid
CA$60K - CA$70K/yr
... for system reliability, security, automation, and documentation. Suggest and implement ... Requirements 2+ years of SysOps,SREor DevOps experience. Strong knowledge in Linux and Windows ...
Site Reliability Specialist
Sorel-tracy, QC · Hybrid
CA$60K - CA$70K/yr
... for system reliability, security, automation, and documentation. Suggest and implement ... Requirements 2+ years of SysOps,SREor DevOps experience. Strong knowledge in Linux and Windows ...
Site Reliability Engineer information
See Quebec salary details
$62.5K - $73K
1% of jobs
$73K - $83.5K
3% of jobs
$83.5K - $94K
5% of jobs
$94K - $104.5K
9% of jobs
$109.8K is the 25th percentile. Wages below this are outliers.
$104.5K - $115K
14% of jobs
$115K - $125.5K
15% of jobs
The median wage is $127.8K / yr.
$125.5K - $136K
15% of jobs
$136K - $146.5K
13% of jobs
$146.9K is the 75th percentile. Wages above this are outliers.
$146.5K - $157K
13% of jobs
$157K - $167.5K
7% of jobs
$167.5K - $178K
5% of jobs
$62.5K
$130.1K
$178K
How much do site reliability engineer jobs pay per year?
Is a site reliability engineer a stressful job?
What is a site reliability engineer?
A site reliability engineer specializes in site reliability engineering, or SRE, a specific branch of operations first pioneered by Google. You are responsible for ensuring that when a website decides to scale a particular feature for various users to access, it does not break the underlying software or website functions. This means you need to use analytical problem-solving skills to determine how to make specific features on a new software release work on top of existing source code.
What are the key skills and qualifications needed to thrive as a site reliability engineer?
What are some of the most common challenges site reliability engineers face when balancing system reliability with rapid software delivery?
What is the difference between Site Reliability Engineer vs DevOps Engineer?
| Aspect | Site Reliability Engineer | DevOps Engineer |
|---|---|---|
| Credentials | Typically requires a computer science degree, certifications like AWS, Google Cloud, or Kubernetes | Similar credentials, often with cloud certifications and scripting skills |
| Work Environment | Focuses on maintaining and improving system reliability, often in large-scale production environments | Works on automation, CI/CD pipelines, and deployment processes across development and operations teams |
| Industry Usage | Common in tech, cloud services, and large-scale enterprise companies | Widely used in software development, cloud, and IT organizations |
Both roles require strong technical skills and cloud knowledge, but SREs focus more on system reliability and uptime, while DevOps engineers emphasize automation and deployment processes. They often collaborate but have distinct primary responsibilities.
What is a site reliability engineer?

Full-time
Medical, Dental, Vision, PTO
Posted 20 days ago
Job description
We are looking for a Site Reliability Engineer (Monetization) who is pragmatic, product-minded, and equally comfortable writing code and running production systems to join our Infrastructure department's SRE team. The best candidate will be someone who thrives in a fast-paced, highly collaborative, and exceptionally dynamic setting and is excited to own the application-level infrastructure and reliability of a high-traffic commerce domain end to end - from deploy pipelines and Kubernetes manifests to SLOs, capacity planning, and production readiness.
Strong Kubernetes, observability, and software engineering skills are essential, along with experience in operating production services in a cloud environment (GCP/GKE or comparable) and partnering closely with product development teams. The ability to hold a dual perspective - understanding both how developers ship features and what infrastructure needs to stay reliable - and to bring the reliability lens into design decisions early will be key to your success in this role.
This is a hybrid embedded role: you remain part of the SRE organization (practices, standards, duty rotation) while being functionally embedded into the Monetization product domain. You'll build long-term working relationships with the domain's engineering teams, own a meaningful share of their application infrastructure execution, and co-author the reliability practices used company-wide.
If you're passionate about making complex distributed systems boringly reliable and love building the commerce and monetization backbone that lets game developers around the world get paid, we would love to hear from you!
ABOUT US
Xsolla is a global commerce company with robust tools and services to help developers solve the inherent challenges of the video game industry. From indie to AAA, companies partner with Xsolla to help them fund, distribute, market, and monetize their games. Grounded in the belief in the future of video games, Xsolla is resolute in the mission to bring opportunities together, and continually make new resources available to creators. Headquartered and incorporated in Los Angeles, California, Xsolla operates as the merchant of record and has helped over 1,500+ game developers to reach more players and grow their businesses around the world. With more paths to profits and ways to win, developers have all the things needed to enjoy the game.
For more information, visit xsolla.com.
- Own the application-level infrastructure of the Monetization domain: Helm charts, Terraform configurations, Kubernetes deployments, runtime configuration, and service-level networking and integrations
- Own the domain's observability: design and implement SLOs/SLIs, monitors, alerts, and dashboards for critical services on Datadog and OpenTelemetry-based tooling
- Help to set up and evolve CI/CD pipelines for domain services (GitLab CI, GitHub Actions), including deploy and rollback automation
- Perform capacity planning and performance tuning ahead of expected load - product launches, sales events, and regional rollouts - including load testing and performance regression investigation
- Run Production Readiness Reviews for new services and major changes; define and enforce what "production-ready" means for the domain
- Support domain incident response: assist with deep investigation of complex incidents, contribute to post-mortems, drive follow-up reliability improvements, and maintain runbooks
- Build domain-specific automation that reduces operational toil: runbook automation, deploy helpers, recurring operational scripts
- Maintain and drive a forward-looking reliability roadmap for the domain together with product engineering leads
- Participate in product team planning, refinements, and architecture reviews, bringing the reliability perspective before design decisions become expensive to change
- Co-author company-wide SLO/SLI, capacity, and operational standards together with the broader SRE team; contribute improvements directly to shared SRE-operated subsystems
- Participate in the SRE duty rotation, supporting developers across the company
- 3+ years of proven SRE, DevOps, or platform engineering experience: on-call or incident response duty, SLO/monitoring ownership, deploy pipeline and infrastructure work for production services
- Software development background: you have built and shipped backend services, not only operated them - comfortable reading application code during an investigation and writing production-quality automation in at least one language (e.g., Go, PHP)
- Hands-on Kubernetes experience: Helm, manifests, deploy strategies, debugging application-level performance and networking issues (GKE or another managed Kubernetes)
- Solid observability practice: building monitors, dashboards, and SLOs/SLIs on a modern platform (Datadog preferred; Prometheus/Grafana experience also relevant), familiarity with OpenTelemetry
- Infrastructure as Code exposure (Terraform/Terragrunt) for collaboration with platform teams
- GCP experience (IAM, networking, managed services)
- Experience building and maintaining CI/CD pipelines (GitLab CI and/or GitHub Actions)
- Programming/scripting proficiency sufficient to build automation and tooling (e.g., Python, Go, or Bash)
- Practical experience with incident response, post-mortems, and driving reliability improvements from incidents
- Strong collaboration and communication skills - this role works embedded with product development teams daily
- Experience in payments, fintech, e-commerce, or gaming - high-traffic transactional systems
- Kubernetes certifications
- Google Cloud Platform certifications
- HashiCorp certifications
We are passionate about fostering a supportive environment for our team, so we prioritize the physical, mental, and emotional well-being of our employees and their families through a comprehensive Benefits Program. This includes medical, dental, and vision, PTO, and a personalized career roadmap for each employee. By investing in professional development through training and educational opportunities, we ensure that our team thrives both personally and professionally. Together, we're not just building a business; we're cultivating a community that values creativity, collaboration, and the transformative power of play.
Equal Employment Opportunity StatementXsolla is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate based on race, color, religion, sex, national origin, age, disability, sexual orientation, gender identity, or any other characteristic protected by law. We consider qualified applicants with criminal histories in accordance with the Fair Chance Act.
Criminal History ConsiderationFor the Site Reliability Engineer (Monetization) position, we will conduct a background check that may include the following:
- Criminal history check
- Employment verification
- Education verification
The background check is relevant to this position because of the following role responsibilities:
- Accessing confidential company data
- Handling infrastructure that processes sensitive financial transactions
- Ensuring compliance with regulatory requirements
Applicants are encouraged to inquire about their rights under the Fair Chance Act. If you have questions regarding our hiring practices, please contact [email protected].
By submitting the following job application form, you consent to Xsolla processing your data for career-related inquiries and potential employment opportunities. We process your data in accordance with this Xsolla Privacy Notice for Job Applicants. Please direct any inquiries regarding your data privacy to [email protected].
About Xsolla
Sourced by ZipRecruiter
Industry
Pc games
Company size
201 - 500 Employees
Headquarters location
Los Angeles, CA, US
Year founded
2005