1

Cloud Operations Engineer Jobs in Toronto, ON (NOW HIRING)

Maintaining, managing, and scaling cloud and container infrastructure for existing and new services, with a strong focus on security, reliability, and operational excellence. * Designing, building ...

Join Capco to help transform enterprise data and analytics through cloud-native engineering and continuous delivery. The Role We're seeking a DevOps Engineer with strong Azure cloud and DevOps ...

Our team is looking for a DevOps engineer with at least 6 years cloud and Kubernetes experience including specialization in ArgoCD. Our project involves building automations for a greenfield managed ...

You will be at the forefront of managing and optimizing our cloud operations for the cloud ... Self-sufficient, works under the supervision of a more senior engineer. * Strong communication ...

Evangelize DevOps culture across the organization, providing training and support to teams on best practices and tooling. To Land This OpportunityYou have 5+ years of experience in DevOps or Cloud ...

DevOps Cloud Engineer

Toronto, ON · On-site

CA$86K - CA$118K/yr

As a DevOps Cloud Engineeryou'llplay a critical role in designing, automating, andoperatingthe cloud infrastructure that powers Nasdaq's global solutions and enables seamless software delivery at ...

DevOps Cloud Engineer

Toronto, ON · Hybrid

CA$86K - CA$118K/yr

As a DevOps CloudEngineeryou'llplay a critical role in designing, automating, andoperatingthe cloud infrastructure that powers Nasdaq's global solutions and enables seamless software delivery at ...

OpenShift Engineer Experience 6-10 Years Required Skills Hands-on experience with Red Hat OpenShift ... Familiarity with Prometheus/Grafana and cloud platforms (AWS/Azure/GCP) is a plus. Key ...

Showing results 21-40

Cloud Operations Engineer information

What is a cloud operations engineer?

A Cloud Operations Engineer is a professional responsible for managing and maintaining an organization's cloud infrastructure and services. Their main duties include deploying, monitoring, and optimizing cloud systems, ensuring security and compliance, and troubleshooting issues related to cloud resources. They work closely with development and IT teams to ensure seamless operation and scalability of cloud-based applications and services. Cloud Operations Engineers are skilled in cloud platforms such as AWS, Azure, or Google Cloud, and often use automation tools to streamline cloud management tasks.

What does a cloud operations engineer do?

Cloud operations engineers use and create cloud-based software and systems to meet their client's needs. As a cloud operations engineer, you collaborate with other cloud operations specialists to design, plan, create, and implement cloud-based software into other operations of a technology company. These applications are designed to streamline operations and make processes more accessible, such as customer support, analysis, or financial reporting. You also work to ensure the plan helps to prevent security breaches and other risks. These positions are usually full-time during regular business hours. Many cloud operations engineers are on the staff of a company's IT department, while others are independent contractors.

What are the key skills and qualifications needed to thrive as a cloud operations engineer, and why are they important?

To thrive as a Cloud Operations Engineer, you need a solid understanding of cloud platforms (such as AWS, Azure, or Google Cloud), scripting, and systems administration, typically supported by a degree in computer science or a related field. Familiarity with infrastructure-as-code tools (like Terraform or CloudFormation), monitoring solutions, CI/CD pipelines, and relevant cloud certifications is highly valued. Strong problem-solving abilities, attention to detail, and effective communication are essential soft skills in this role. These skills are crucial for ensuring reliable, secure, and efficient cloud operations that support business objectives.

What are some common challenges faced by cloud operations engineers, and how can they be addressed?

Cloud Operations Engineers often encounter challenges such as ensuring system uptime, managing security across distributed environments, and rapidly responding to incidents. To address these, it's important to implement robust monitoring tools, automate routine tasks, and maintain clear documentation. Collaboration with development, security, and support teams is also key—regular communication helps to proactively identify and resolve potential issues, making the role both dynamic and highly collaborative.

Are cloud operations engineers still in demand?

Cloud operations engineers are currently in high demand due to the ongoing growth of cloud computing and digital transformation across industries. They are needed to manage cloud infrastructure, optimize performance, and ensure security, often requiring skills in platforms like AWS, Azure, or Google Cloud, along with certifications such as AWS Certified Solutions Architect or CompTIA Cloud+. The role is expected to remain vital as organizations continue migrating to cloud environments.

Is a cloud operations engineer a high paying job?

Cloud operations engineers typically earn above-average salaries due to their specialized skills in managing cloud infrastructure, automation, and security. Compensation varies based on experience, certifications, and location, but it is generally considered a well-paying role within the IT industry.

What are the most commonly searched types of Cloud Operations Engineer jobs in Toronto, ON?

The most popular types of Cloud Operations Engineer jobs in Toronto, ON are:

What are popular job titles related to Cloud Operations Engineer jobs in Toronto, ON?

For Cloud Operations Engineer jobs in Toronto, ON, the most frequently searched job titles are:

What job categories do people searching Cloud Operations Engineer jobs in Toronto, ON look for?

The top searched job categories for Cloud Operations Engineer jobs in Toronto, ON are:

Infographic showing various Cloud Operations Engineer job openings in Toronto, ON as of August 2026, with employment types broken down into 85% Full Time, 12% Part Time, 1% Temporary, 1% Contract, and 1% Nights. Highlights an 92% Physical, 4% Hybrid, and 4% Remote job distribution.

Senior DevOps Engineer, Agentic Automation & Live Games

Big Viking Games

Toronto, ON • On-site

CA$140K - CA$180K/yr

Full-time

Medical, Dental, Vision, Retirement

Posted 12 days ago


Key responsibilities

  • Build agentic automation tools that perform diagnosis, remediation, provisioning, and routine maintenance tasks.

  • Monitor, maintain, and improve cloud infrastructure supporting live games and related systems.

  • Improve observability, reliability, and security of the production environment through logging, metrics, alerting, and operational visibility.


Job description

Senior DevOps Engineer, Agentic Automation & Live Games

Toronto, ON
Hybrid, 3 days per week in office
Full-time
Department: Engineering
Reports to: Engineering Leadership
Compensation range: CAD $140,000 to $180,000

About Big Viking Games

Big Viking Games is a Canadian gaming company focused on building, operating, and growing long-standing online game communities. Our games have entertained players for years, supported by loyal audiences, live operations, evolving content systems, product innovation, and deep player-driven economies.

Our flagship titles, YoWorld and FishWorld, have served millions of players over their lifetime. These are enduring live-service virtual worlds with rich in-game economies, virtual goods, social interaction, and long-term player engagement at their core.

We are entering a new phase of modernization and growth, with a focus on stronger infrastructure, better automation, practical AI adoption, improved reliability, stronger security practices, and scalable systems that help our games and teams perform at a higher level.

About the Role

Big Viking Games is hiring a Senior DevOps Engineer, Agentic Automation & Live Games to modernize the infrastructure behind our live-service games and help reinvent how those systems are operated.

This is a hands-on senior infrastructure role for someone fluent in modern cloud operations: AWS, containers, Infrastructure as Code, CI/CD, observability, production reliability, incident response, security, and automation.

But this is not a traditional DevOps role.

We are looking for someone who can build automation that does real work, not just scripts that run or dashboards that summarize. The right person has started using agentic coding tools, tool-calling systems, API integrations, workflow automation, or MCP-style tooling to safely diagnose, provision, remediate, monitor, or maintain infrastructure.

Our games run on mature systems with real players, real revenue, real constraints, and real consequences. There is meaningful room to automate how they are operated, but the work must be done with discipline. Uptime, data integrity, least-privilege access, rollback paths, auditability, and production safety matter.

The defining trait for this role is self-direction. Given a backlog, you improve how the work gets done. Left to your own judgment, you find repetitive operational work nobody has flagged, decide what is worth automating, and build it safely.

This is a hybrid role based in Toronto, with an expectation of working in office three days per week. Live-service games require operational awareness outside regular business hours, including periodic on-call and incident response availability.

What You’ll DoBuild Agentic Automation for Infrastructure Work
  • Design, build, and operate agentic tooling that performs real DevOps work, including diagnosis, remediation, provisioning, routine maintenance, and operational follow-up.
  • Build and maintain the integration layer that lets tooling act safely on our systems, including API integrations, MCP-style servers, webhook-driven orchestration, serverless functions, and permission-controlled automation.
  • Convert manual runbooks, SOPs, recurring maintenance tasks, and repetitive operational chores into automation that can run unattended where appropriate.
  • Establish guardrails that make automated action against production defensible, including least-privilege scopes, dry-run modes, approval paths, logging, audit trails, rollback plans, and clear escalation rules.
  • Use agentic coding tools to accelerate infrastructure work, including IaC authoring, migration scripts, incident analysis, log analysis, documentation, and operational troubleshooting.
  • Help improve how the wider engineering team uses automation and AI-assisted tooling safely and effectively.
Own and Modernize Live Production Infrastructure
  • Monitor, maintain, and improve cloud infrastructure across AWS, Netlify, Vercel, and related platforms supporting our live games, data systems, internal tools, and operational workflows.
  • Drive infrastructure modernization while maintaining uptime for live games with active player communities.
  • Implement and maintain Infrastructure as Code using Terraform, CloudFormation, CDK, Pulumi, or similar tools.
  • Improve CI/CD pipelines, release workflows, deployment reliability, and environment management so teams can ship safely and frequently.
  • Operate containerized workloads and GitOps-based deployment patterns.
  • Improve local development, build, test, deploy, and production support workflows.
  • Help manage cloud usage, resource tagging, environment efficiency, and infrastructure cost discipline.
Reliability, Observability, and Security
  • Rebuild and improve observability across the stack, including logging, metrics, alerting, dashboards, and operational visibility.
  • Pay particular attention to early detection of silent failures, pipeline failures, data freshness issues, and degraded production behavior.
  • Maintain and monitor data pipelines between game source databases, Snowflake, and downstream analytics and reporting systems.
  • Own secrets and credential lifecycle management across platforms, including API key rotation, access controls, environment variable governance, and least-privilege practices.
  • Support incident response, root cause analysis, remediation planning, and post-incident improvements.
  • Automate the parts of incident response and remediation that repeat.
  • Create documentation, runbooks, and SOPs that are executable wherever possible, not just descriptive.
  • Improve backup, restore, disaster recovery, access review, and production-readiness practices.

Requirements

What You BringExperience
  • 7+ years of experience in DevOps, infrastructure engineering, cloud engineering, site reliability engineering, platform engineering, or a similar role.
  • Strong hands-on experience designing, maintaining, and improving production infrastructure.
  • Experience supporting live production systems where uptime, reliability, data integrity, and performance matter.
  • Experience working with mature or legacy systems that predate modern cloud-native patterns.
  • 1+ year building with agentic coding tools, tool-calling systems, workflow automation, or infrastructure automation that takes real action.
  • Experience with systems such as agents wired into pipelines, MCP-style integrations, automated diagnosis, automated remediation, provisioning workflows, or production-safe infrastructure tooling.
  • Experience identifying repetitive operational work and turning it into reliable automation.
Core Technical Skills
  • Strong hands-on experience with AWS or similar cloud platforms.
  • Proficiency with Infrastructure as Code tools such as Terraform, CloudFormation, CDK, Pulumi, or similar, including shared state and team-based workflows.
  • Experience with containerized applications, especially Docker.
  • Experience with container orchestration such as Kubernetes, ECS, EKS, or similar, including debugging real production issues.
  • Experience with GitOps and declarative deployment tools such as ArgoCD, Flux, or equivalent.
  • Experience with CI/CD tooling, version control, deployment automation, and modern release workflows.
  • Strong understanding of Linux systems, networking, cloud security, monitoring, logging, and operational troubleshooting.
  • Experience establishing observability, including choosing tooling, defining alerts, tuning alert noise, and creating useful dashboards.
  • Experience with relational databases such as MariaDB, MySQL, or Postgres, including replication, backup, verified restore, and schema changes against systems that stay online.
  • Security-aware mindset with practical experience in secrets management, credential rotation, access control, vulnerability reduction, and least-privilege practices.
How You Work
  • Self-starting. You identify the work rather than waiting for it to be assigned, and you can explain why one problem matters more than another.
  • Proactive about toil. You notice repetitive work and treat it as a defect to be engineered away, not a cost of doing business.
  • Inventive but pragmatic. You reach for novel approaches where they genuinely help and recognize when a simple script is the better answer.
  • Strong problem-solving skills and the ability to investigate complex infrastructure or production issues, including silent failures and data pipeline outages.
  • Comfortable working directly with engineers to improve build, deploy, and operational workflows.
  • Practical ownership mindset with the ability to prioritize, execute, and close loops.
  • Strong communication with technical and non-technical stakeholders.
  • Calm under pressure during incidents, escalations, and ambiguous production issues.
Nice to Have
  • Experience supporting live games, virtual worlds, multiplayer systems, or other real-time online products.
  • Experience with GitHub Actions or similar CI/CD platforms.
  • Experience with Datadog, Grafana, Prometheus, CloudWatch, ELK, OpenTelemetry, or similar observability tooling.
  • Experience with Redis, Memcached, queues, workers, or event-driven systems.
  • Experience with Snowflake, data warehouse connectivity, ETL monitoring, or data pipeline reliability.
  • Experience with serverless platforms such as Netlify Functions, Vercel, AWS Lambda, or similar.
  • Experience operating multi-platform hosting environments.
  • Experience with disaster recovery, backup strategies, incident management, load testing, and performance tuning.
  • Experience improving cloud cost management, tagging, resource optimization, or infrastructure governance.
  • Experience evaluating where automation should not be applied, including cases where human approval or human judgment remained necessary.
  • Experience working in small, high-leverage engineering teams where infrastructure ownership is broad and hands-on.
Ideal Candidate Profile

The ideal candidate is a senior infrastructure engineer who keeps live systems stable while systematically removing the manual work involved in keeping them that way.

They are not a tool collector. They understand uptime, production risk, cloud cost, security, release quality, operational discipline, and the failure modes of automation. They know that automation acting on production needs guardrails, observability, rollback paths, and human approval where risk demands it.

They work independently. Given a mature live game and a broad remit, they can identify what matters, sequence it sensibly, and improve systems incrementally without disrupting what is already working.

They are comfortable being the person who decides what gets automated next, what should remain manual, and what needs stronger controls before it can safely run unattended.

This role suits someone who wants deep ownership of production infrastructure and a real mandate to modernize how live-service games are operated.

Benefits

Compensation

The expected compensation range for this role is CAD $140,000 to $180,000, depending on experience, technical depth, infrastructure ownership, agentic automation experience, and overall fit.

Benefits
  • Group Retirement Savings Plan matching and participation.
  • Comprehensive benefits package, including health, dental, and vision coverage.
  • Health and Wellness spending account.
  • Generous time off policies.
  • Opportunity to support long-running live-service games with established player communities.
  • Deep ownership of cloud modernization, DevOps automation, security improvement, and agentic infrastructure workflows.
  • A high-impact role with meaningful ownership over reliability, performance, and engineering operations.
Hiring Process and AI Disclosure

This posting is for an existing vacancy.

Big Viking Games may use AI-assisted tools at some steps of the recruiting process, including application review support, candidate research, interview preparation, scheduling support, and workflow administration.

AI does not make final hiring decisions. Hiring decisions are made by people.

Every interviewed candidate will be informed of their status within 45 days of their final interview.

Accessibility and Accommodation

Big Viking Games is committed to creating an inclusive and accessible environment for all candidates. We welcome applications from individuals of all abilities and will provide accommodations throughout the hiring process as needed.

If you require accommodation during the hiring process, please contact hr@bigvikinggames.com so we can work with you to support your needs.

Application Note

When you apply, tell us about infrastructure automation you have built that performed real operational work.

We are especially interested in what it did, what systems it touched, how you made it safe, what guardrails you built, how it failed, and what manual work it eliminated.