1

Infrastructure Monitoring Jobs in Alberta (NOW HIRING)

Senior Java Developer

Calgary, AB ยท On-site +1

$91K - $120K/yr

Monitor and troubleshoot system performance using tools like Prometheus, Graphite and CloudWatch. * Apply Terraform to automate and simplify infrastructure provisioning and configuration management.

Senior Java Developer

Calgary, AB ยท On-site +1

$91K - $120K/yr

Monitor and troubleshoot system performance using tools like Prometheus, Graphite and CloudWatch. * Apply Terraform to automate and simplify infrastructure provisioning and configuration management.

Showing results 21-40

Infrastructure Monitoring information

See Alberta salary details

$11

$60

$103

How much do infrastructure monitoring jobs pay per hour?

As of Aug 6, 2026, the average hourly pay for infrastructure monitoring in Alberta is $60.15, according to ZipRecruiter salary data. Most workers in this role earn between $35.82 and $89.90 per hour, depending on experience, location, and employer.

What are the key skills and qualifications needed to thrive in infrastructure monitoring?

To thrive in Infrastructure Monitoring, you need strong analytical skills, a foundational understanding of IT infrastructure, and experience with system administration or network management. Familiarity with monitoring tools like Nagios, Zabbix, SolarWinds, and relevant certifications such as CompTIA Network+ or Microsoft Certified: Azure Administrator are highly beneficial. Attention to detail, problem-solving abilities, and effective communication help professionals excel in monitoring and responding to system issues quickly. These skills and qualities are essential to ensuring uptime, performance, and security across complex IT environments.

What is an infrastructure monitoring?

An Infrastructure Monitoring job involves overseeing IT systems, networks, and servers to ensure optimal performance and availability. Professionals in this role use monitoring tools to detect issues such as downtime, performance bottlenecks, or security threats. They analyze logs, set up alerts, and take proactive measures to maintain system health. The goal is to minimize disruptions, improve efficiency, and support business continuity.

What are some common challenges faced in infrastructure monitoring?

Professionals in Infrastructure Monitoring often deal with the challenges of identifying and responding to system anomalies in real-time, troubleshooting issues across diverse technologies, and managing high volumes of alerts. The work environment is typically fast-paced, and requires carefully prioritizing incidents to ensure minimal downtime and optimal system performance. Collaboration with network engineers, DevOps teams, and support staff is a regular part of the role. Successfully addressing these challenges requires a mix of technical expertise, adaptability, and strong communication skills, making it a dynamic and rewarding career path for IT specialists.

What are popular job titles related to Infrastructure Monitoring jobs in Alberta? For Infrastructure Monitoring jobs in Alberta, the most frequently searched job titles are:
What job categories do people searching Infrastructure Monitoring jobs in Alberta look for? The top searched job categories for Infrastructure Monitoring jobs in Alberta are:
Infographic showing various Infrastructure Monitoring job openings in Alberta as of July 2026, with employment types broken down into 100% Full Time. Highlights an 100% In-person job distribution, with an average salary of $125,117 per year, or $60.2 per hour.

Software Development Lead- Platform

Krux Analytics

Calgary, AB โ€ข On-site

Full-time

Posted 20 days ago


Job description

Purpose

This role is designed to bring dedicated technical leadership to our platform team while also owning production reliability and quality. Leading a cross-functional platform team, including DevOps specialists, senior data developers/architects, and QA automation, while simultaneously owning the end-to-end production health of our SaaS product. You will be the bridge between engineering, product, and customer success, driving systematic strategies to reduce bugs, improve reliability, and elevate the overall quality of our platform.

Accountabilities & Expected Outcomes

Platform Team Leadership & People Management

  • Lead, mentor, and grow a cross-functional platform team of DevOps specialists, Power BI/data platform architects, and a QA automation engineer.
  • Set clear goals, conduct regular 1:1s, and support the professional development of each team member.
  • Foster a collaborative, high-trust team culture aligned with Team Topologies principles — acting as an enabling team that reduces cognitive load on stream-aligned teams.
  • Manage team capacity, prioritization, and delivery of platform initiatives.
  • Hire and onboard new platform team members as the organization scales.


Production Reliability & Bug Management

  • Own the end-to-end production monitoring strategy, including tooling, alerting, and observability frameworks.
  • Lead bug triage processes, assess severity, coordinate investigation, assign ownership across stream-aligned teams, and track resolution.
  • Drive root cause analysis (RCA) for recurring issues and translate findings into systemic improvement strategies.
  • Define and track reliability metrics (e.g. MTTR, bug escape rate, defect density) and report on trends to engineering and product leadership.
  • Build and continuously improve processes for bug intake, prioritization, and communication to internal stakeholders.

Stakeholder Communication & Cross-Team Coordination

  • Act as the primary technology point of contact for production issues, communicating clearly and proactively with Product and Customer Success teams.
  • Provide regular status updates on open bugs, reliability trends, and resolution timelines.
  • Coordinate with stream-aligned teams to ensure production issues are addressed efficiently without disrupting their delivery flow.
  • Translate technical complexity into clear, business-relevant language for non-technical stakeholders.


Platform Engineering & Enablement

  • Oversee the design, development, operation, and documentation of shared engineering infrastructure including CI/CD pipelines, cloud infrastructure, monitoring tools, and the internal Power BI data platform.
  • Ensure the platform team delivers high-quality, self-service capabilities that reduce friction for product engineering teams.
  • Champion engineering standards: security baselines, observability practices, testing frameworks, and deployment best practices.
  • Identify and evaluate build-vs-buy decisions for platform tooling; manage relevant vendor relationships.
  • Collaborate with stream-aligned teams to understand their needs and shape the platform roadmap accordingly.


Quality Engineering Strategy

  • Leverage the QA automation engineer to build scalable test frameworks that integrate into CI/CD pipelines.
  • Drive a shift-left quality mindset across engineering teams, embedding quality practices earlier in the development lifecycle.
  • Use production data and bug patterns to inform test coverage improvements and preventive quality strategies.


Leadership and People

  • Sharing and cascading the IMDEX vision, strategy and objectives to own team.
  • Team members understand the IMDEX vision and team objectives with clear role clarity.
  • Ensuring staff are engaged and competent to perform their tasks and compliant with all group policies, procedures and regional regulations.
  • Raising difficult issues to ensure that they are addressed and leading team members to follow through on difficult actions or initiatives.
  • Willingly passes on skills, sharing job knowledge and training others.
  • Assigning appropriate support and resources to personnel to maximize success.
  • Setting and providing clear direction and expectations. Effectively delegating and removing obstacles to get work done. Setting stretch goals and opportunities to drive higher performance.
  • Team employee engagement levels increase annually through investment in people.
  • Recruiting talent based on capability requirements.
  • Performance and Development Reviews (Achievement and Development Reviews ADRs) on direct reports completed, ensuring cascade to full region.
  • Poor performance issues addressed in a timely manner.
  • Supporting team members to create and action development objectives.
  • Learning from setbacks, keeping calm under pressure and is sensitive to the needs of others during a time of crisis.


Health, Safety and Environment

  • Utilizing resources to establish, implement, maintain, and improve the Quality, Health, Safety and Environment (QHSE) management system.
  • Ensuring risk management activities are conducted within area of responsibility, including the identification and implementation of any site-specific measures required to eliminate or reduce risk in their area.
  • Ensuring employees are provided with the necessary personal protective equipment (PPE), instruction, information, training, and supervision to enable work to be carried out safely.
  • In consultation with the QHSE Representative, developing and implementing an injury management program for injured employees.
  • Completion of hazard and aspect identification and elimination/reduction programmes for team.
  • Effective incident reporting.
  • Regular workplace inspections conducted.
  • Safety meetings conducted.
  • Site activities comply with workplace health, safety, and environmental legislation and the IMDEX QHSE management system.
  • Employees and QHSE Representatives are consulted with regarding any QHSE issues and performance.
  • Quality alert system is utilized to initiate and respond to QHSE issues.


Risk and Compliance

  • Leading the collaborative identification of risks and compliance requirements relevant to the team.
  • Promoting the reporting of potential non-compliance or misconduct through appropriate channels.
  • Leading by example to promote positive risk and compliance behaviors to the team and to colleagues.
  • Ensuring the compliance of team members with IMDEX risk and compliance policies, local laws and other relevant standards.
  • Applicable risk management and regulatory compliance processes are implemented by the team.
  • Behaviors that do not align with the IMDEX Code of Conduct are challenged.


Qualifications, Skills & Experience

  • Bachelor’s Degree (Computer Science, Technology, Engineering, or related field)
  • 5+ years of experience in DevOps, platform engineering, SRE, or a related technical discipline such as software development or technical QA.
  • 5+ years in a people management or technical lead role, with demonstrated success in growing engineering talent.
  • Hands-on experience with CI/CD pipelines, cloud infrastructure (Azure, AWS, or GCP), and monitoring/observability tooling.
  • Proven track record of owning production reliability, driving incident response, and implementing systematic quality improvements.
  • Strong communication skills, able to engage effectively with both engineers and non-technical business stakeholders.
  • Familiarity with agile delivery practices and experience working within product engineering organizations.
  • Experience in a SaaS product company with production quality responsibilities.
  • Demonstrated ability to manage multiple competing priorities and translate strategy into actionable plans.


Desirable

  • Product reliability
  • Quality Assurance
  • Verbal and written communication
  • Mentoring
  • Leadership
  • Strategic thinking
  • Conflict Resolution
  • Goal setting
  • Analytical
  • Team Building
  • Critical thinking
  • Adaptability
  • Time management
  • Project Management
  • Problem Solving
  • Technical Proficiency
  • Empathy