1

Site Reliability Engineer Jobs in Quebec (NOW HIRING)

About the Role We are looking for a Site Reliability Engineer to join our Network and Security Operations Center (NOC), a team at the heart of platform reliability for mission-critical SaaS ...

About the Role We are looking for a Site Reliability Engineer to join our Network and Security Operations Center (NOC), a team at the heart of platform reliability for mission-critical SaaS ...

The SRE will function as part of a special investigations unit that empowers and enables Applicative Support, Infrastructure Support, and the Incident Management team-coaching, guiding, and leading ...

Nous sommes a la recherche d'un(e) ingenieur(e) en fiabilite des sites (Site Reliability Engineer - SRE) competent(e) pour concevoir, exploiter et ameliorer des systemes de production hautement ...

We're seeking someone to join our team as a Site Reliability Engineering Specialist to support the reliability, performance, and operational stability of business-critical applications, while helping ...

To do that we are eager to add a highly skilled DevOps / SRE Engineer Engineer to our incredible team. This is a senior role working alongside the current backend and frontend software engineering ...

As a Site Reliability Specialist at Ubisoft Montréal, you will join the IT Games and Studios team ... You'll collaborate with developers, cloud specialists, and infrastructure teams to build resilient ...

The COO/GTE/EPL/SRE team has members in Paris, Bangalore, and Montreal and is responsible for the production, security, performance, and scalability of all capabilities provided by EPL. WHAT WILL BE ...

next page

Showing results 1-20

Site Reliability Engineer information

See Quebec salary details

$62.5K

$130.1K

$178K

How much do site reliability engineer jobs pay per year?

As of Aug 6, 2026, the average yearly pay for site reliability engineer in Quebec is $130,104.00, according to ZipRecruiter salary data. Most workers in this role earn between $109,500.00 and $149,500.00 per year, depending on experience, location, and employer.

Is a site reliability engineer a stressful job?

A site reliability engineer (SRE) role can be stressful due to the responsibility of maintaining system uptime, handling incidents, and ensuring reliability under tight deadlines. The job often involves on-call duties, troubleshooting complex issues, and working with automation tools, which can contribute to work-related stress but also offers opportunities for skill development and problem-solving.

What is a site reliability engineer?

A site reliability engineer specializes in site reliability engineering, or SRE, a specific branch of operations first pioneered by Google. You are responsible for ensuring that when a website decides to scale a particular feature for various users to access, it does not break the underlying software or website functions. This means you need to use analytical problem-solving skills to determine how to make specific features on a new software release work on top of existing source code.

What are the key skills and qualifications needed to thrive as a site reliability engineer?

To thrive as a Site Reliability Engineer, you need a strong background in computer science, systems administration, and software engineering, often supported by a degree in a technical field. Familiarity with cloud platforms (like AWS or GCP), container orchestration (such as Kubernetes), infrastructure as code (Terraform or Ansible), and monitoring tools (Prometheus, Grafana) is typically expected. Strong problem-solving skills, effective communication, and a proactive mindset help SREs excel at incident management and cross-functional collaboration. These skills are crucial for maintaining system reliability, minimizing downtime, and driving continuous improvement in complex technical environments.

What are some of the most common challenges site reliability engineers face when balancing system reliability with rapid software delivery?

Site Reliability Engineers (SREs) often navigate the challenge of maintaining highly reliable systems while supporting fast-paced software releases. This involves managing incidents, automating processes to reduce manual toil, and working closely with development teams to embed reliability into the software development lifecycle. SREs must carefully prioritize their efforts between proactive improvements and urgent, reactive fire-fighting. Effective communication and collaboration with both operations and development teams are crucial to ensuring service uptime without slowing down innovation.

What is the difference between Site Reliability Engineer vs DevOps Engineer?

AspectSite Reliability EngineerDevOps Engineer
CredentialsTypically requires a computer science degree, certifications like AWS, Google Cloud, or KubernetesSimilar credentials, often with cloud certifications and scripting skills
Work EnvironmentFocuses on maintaining and improving system reliability, often in large-scale production environmentsWorks on automation, CI/CD pipelines, and deployment processes across development and operations teams
Industry UsageCommon in tech, cloud services, and large-scale enterprise companiesWidely used in software development, cloud, and IT organizations

Both roles require strong technical skills and cloud knowledge, but SREs focus more on system reliability and uptime, while DevOps engineers emphasize automation and deployment processes. They often collaborate but have distinct primary responsibilities.

What is a site reliability engineer?

A Site Reliability Engineer (SRE) is a professional who applies software engineering principles to infrastructure and operations problems. Their primary goal is to create scalable and highly reliable software systems, often bridging the gap between development and IT operations. SREs automate tasks, monitor system health, respond to incidents, and work to improve system reliability and performance. They also help define service level objectives (SLOs) and ensure systems meet customer expectations for uptime and availability.
What are the most commonly searched types of Site Reliability Engineer jobs in Quebec? The most popular types of Site Reliability Engineer jobs in Quebec are:
What are popular job titles related to Site Reliability Engineer jobs in Quebec? For Site Reliability Engineer jobs in Quebec, the most frequently searched job titles are:
What job categories do people searching Site Reliability Engineer jobs in Quebec look for? The top searched job categories for Site Reliability Engineer jobs in Quebec are:
What are popular job titles related to Site Reliability Engineer jobs in QC? For Site Reliability Engineer jobs in QC, the most frequently searched job titles are:
Infographic showing various Site Reliability Engineer job openings in Quebec as of July 2026, with employment types broken down into 94% Full Time, 3% Part Time, and 3% Contract. Highlights an 85% Physical, 5% Hybrid, and 10% Remote job distribution, with an average salary of $130,104 per year, or $62.5 per hour.

Site Reliability Engineer (Monetization)

Xsolla

Montreal, QC • On-site

Full-time

Medical, Dental, Vision, PTO

Posted 20 days ago


Job description

ABOUT YOU 

We are looking for a Site Reliability Engineer (Monetization) who is pragmatic, product-minded, and equally comfortable writing code and running production systems to join our Infrastructure department's SRE team. The best candidate will be someone who thrives in a fast-paced, highly collaborative, and exceptionally dynamic setting and is excited to own the application-level infrastructure and reliability of a high-traffic commerce domain end to end - from deploy pipelines and Kubernetes manifests to SLOs, capacity planning, and production readiness.

Strong Kubernetes, observability, and software engineering skills are essential, along with experience in operating production services in a cloud environment (GCP/GKE or comparable) and partnering closely with product development teams. The ability to hold a dual perspective - understanding both how developers ship features and what infrastructure needs to stay reliable - and to bring the reliability lens into design decisions early will be key to your success in this role.

This is a hybrid embedded role: you remain part of the SRE organization (practices, standards, duty rotation) while being functionally embedded into the Monetization product domain. You'll build long-term working relationships with the domain's engineering teams, own a meaningful share of their application infrastructure execution, and co-author the reliability practices used company-wide.

If you're passionate about making complex distributed systems boringly reliable and love building the commerce and monetization backbone that lets game developers around the world get paid, we would love to hear from you!

ABOUT US

Xsolla is a global commerce company with robust tools and services to help developers solve the inherent challenges of the video game industry. From indie to AAA, companies partner with Xsolla to help them fund, distribute, market, and monetize their games. Grounded in the belief in the future of video games, Xsolla is resolute in the mission to bring opportunities together, and continually make new resources available to creators. Headquartered and incorporated in Los Angeles, California, Xsolla operates as the merchant of record and has helped over 1,500+ game developers to reach more players and grow their businesses around the world. With more paths to profits and ways to win, developers have all the things needed to enjoy the game.

For more information, visit xsolla.com.

Responsibilities
  • Own the application-level infrastructure of the Monetization domain: Helm charts, Terraform configurations, Kubernetes deployments, runtime configuration, and service-level networking and integrations
  • Own the domain's observability: design and implement SLOs/SLIs, monitors, alerts, and dashboards for critical services on Datadog and OpenTelemetry-based tooling
  • Help to set up and evolve CI/CD pipelines for domain services (GitLab CI, GitHub Actions), including deploy and rollback automation
  • Perform capacity planning and performance tuning ahead of expected load - product launches, sales events, and regional rollouts - including load testing and performance regression investigation
  • Run Production Readiness Reviews for new services and major changes; define and enforce what "production-ready" means for the domain
  • Support domain incident response: assist with deep investigation of complex incidents, contribute to post-mortems, drive follow-up reliability improvements, and maintain runbooks
  • Build domain-specific automation that reduces operational toil: runbook automation, deploy helpers, recurring operational scripts
  • Maintain and drive a forward-looking reliability roadmap for the domain together with product engineering leads
  • Participate in product team planning, refinements, and architecture reviews, bringing the reliability perspective before design decisions become expensive to change
  • Co-author company-wide SLO/SLI, capacity, and operational standards together with the broader SRE team; contribute improvements directly to shared SRE-operated subsystems
  • Participate in the SRE duty rotation, supporting developers across the company
Qualifications & Skills
  • 3+ years of proven SRE, DevOps, or platform engineering experience: on-call or incident response duty, SLO/monitoring ownership, deploy pipeline and infrastructure work for production services
  • Software development background: you have built and shipped backend services, not only operated them - comfortable reading application code during an investigation and writing production-quality automation in at least one language (e.g., Go, PHP)
  • Hands-on Kubernetes experience: Helm, manifests, deploy strategies, debugging application-level performance and networking issues (GKE or another managed Kubernetes)
  • Solid observability practice: building monitors, dashboards, and SLOs/SLIs on a modern platform (Datadog preferred; Prometheus/Grafana experience also relevant), familiarity with OpenTelemetry
  • Infrastructure as Code exposure (Terraform/Terragrunt) for collaboration with platform teams
  • GCP experience (IAM, networking, managed services)
  • Experience building and maintaining CI/CD pipelines (GitLab CI and/or GitHub Actions)
  • Programming/scripting proficiency sufficient to build automation and tooling (e.g., Python, Go, or Bash)
  • Practical experience with incident response, post-mortems, and driving reliability improvements from incidents
  • Strong collaboration and communication skills - this role works embedded with product development teams daily
  • Experience in payments, fintech, e-commerce, or gaming - high-traffic transactional systems
Nice to Have:
  • Kubernetes certifications
  • Google Cloud Platform certifications
  • HashiCorp certifications
$120,000 - $160,000 a year
Salary varies depending on experience level and location.
Benefits

We are passionate about fostering a supportive environment for our team, so we prioritize the physical, mental, and emotional well-being of our employees and their families through a comprehensive Benefits Program. This includes medical, dental, and vision, PTO, and a personalized career roadmap for each employee. By investing in professional development through training and educational opportunities, we ensure that our team thrives both personally and professionally. Together, we're not just building a business; we're cultivating a community that values creativity, collaboration, and the transformative power of play.

Equal Employment Opportunity Statement

Xsolla is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate based on race, color, religion, sex, national origin, age, disability, sexual orientation, gender identity, or any other characteristic protected by law. We consider qualified applicants with criminal histories in accordance with the Fair Chance Act.

Criminal History Consideration

For the Site Reliability Engineer (Monetization) position, we will conduct a background check that may include the following:

  • Criminal history check
  • Employment verification
  • Education verification
Relevance to Job Responsibilities

The background check is relevant to this position because of the following role responsibilities:

  • Accessing confidential company data
  • Handling infrastructure that processes sensitive financial transactions
  • Ensuring compliance with regulatory requirements
Rights Under the Fair Chance Act

Applicants are encouraged to inquire about their rights under the Fair Chance Act. If you have questions regarding our hiring practices, please contact [email protected].

By submitting the following job application form, you consent to Xsolla processing your data for career-related inquiries and potential employment opportunities. We process your data in accordance with this Xsolla Privacy Notice for Job Applicants. Please direct any inquiries regarding your data privacy to [email protected].

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
apply for this job