1

Infiniband Jobs in Ontario (NOW HIRING)

CA$130K - CA$175K/yr

Our software runs on production GPU and CPU clusters across Ethernet, RDMA/RoCE, and InfiniBand networks. Testing here means standing up real multi-node clusters, running real workloads against them ...

Site Reliability Engineer

Toronto, ON · On-site +1

CA$125K - CA$250K/yr

Networking, including firewalls, switching, routing, ASN/BGP configuration, or InfiniBand * Cluster and systems allocation with Kubernetes, SLURM, MAAS, or similar platforms * Distributed storage ...

Choose appropriate CPUs, GPUs, interconnects (e.g., InfiniBand), and storage * Configure Slurm, PBS, or OpenHPC job schedulers * Cluster Deployment & Maintenance * Install and manage Linux-based ...

CA$150K - CA$230K/yr

High-performance networking (RDMA, InfiniBand) * ML framework or runtime internals * Cluster scheduling or orchestration systems We believe strong systems engineers pick up domain-specific tools ...

Proven success managing multi-terabit interconnect products or high-speed networking technologies (Ethernet, PCIe, InfiniBand, etc.). * Excellent communication, presentation, and negotiation skills ...

Familiarity with high-performance GPU infrastructure (e.g., NVIDIA H100/H200/B200, InfiniBand networking, parallel file systems) Wondering if you're a good fit? We believe in investing in our people ...

Familiarity with high-performance GPU infrastructure (e.g., NVIDIA H100/H200/B200, InfiniBand networking, parallel file systems) Wondering if you're a good fit? We believe in investing in our people ...

Infiniband information

What is InfiniBand?

Infiniband is a high-speed, low-latency networking technology commonly used in data centers and high-performance computing environments. It is designed to connect servers, storage systems, and network devices, providing much faster data transfer rates than traditional Ethernet. Infiniband supports scalable bandwidth and efficient communication, which makes it ideal for applications requiring rapid data movement, such as scientific simulations and large-scale database transactions. Its architecture also supports remote direct memory access (RDMA), which further reduces latency and CPU overhead.

What are the typical responsibilities of an InfiniBand network engineer in a data center environment?

InfiniBand network engineers are primarily responsible for designing, deploying, and maintaining high-performance InfiniBand fabrics that connect servers and storage systems in data centers, especially in HPC (High-Performance Computing) environments. Their daily tasks include monitoring network performance, troubleshooting connectivity or latency issues, and performing firmware and driver updates on InfiniBand switches and host adapters. They also collaborate closely with system administrators and application teams to optimize throughput and ensure reliable, low-latency communication. Additionally, InfiniBand engineers often participate in capacity planning and help scale the network infrastructure to meet growing computational demands.

What are the key skills and qualifications needed to thrive as an InfiniBand network engineer, and why are they important?

To thrive as an InfiniBand Network Engineer, you need a strong background in computer networking, Linux system administration, and high-performance computing (HPC) environments, often supported by a degree in computer science or related field. Familiarity with InfiniBand architecture, experience with tools like OpenFabrics Enterprise Distribution (OFED), and certifications such as CompTIA Network+ are valuable. Strong problem-solving skills, attention to detail, and effective communication are crucial soft skills for this role. These abilities are essential for ensuring efficient, reliable InfiniBand network performance in complex HPC or data center environments.

What is the difference between Infiniband vs Ethernet Network Engineer?

AspectInfinibandEthernet Network Engineer
Required CredentialsNetworking certifications, Cisco, Cisco CCNA, CCNPNetworking certifications, Cisco, CCNA, CCNP
Work EnvironmentData centers, high-performance computing environmentsCorporate networks, data centers, enterprise environments
Industry UsageHigh-performance computing, research institutionsBusiness, telecommunications, enterprise IT
Common Search/ComparisonYesYes

Infiniband and Ethernet Network Engineers both work with network infrastructure, but Infiniband specializes in high-speed, low-latency connections used in data centers and HPC environments. Ethernet Network Engineers focus on standard Ethernet networks used across various industries. While their certifications and skills overlap, their work environments and applications differ significantly.

What job categories do people searching Infiniband jobs in Ontario look for?

The top searched job categories for Infiniband jobs in Ontario are:

Infographic showing various Infiniband job openings in Ontario as of August 2026, with employment types broken down into 98% Full Time, and 2% Contract. Highlights an 81% Physical, 4% Hybrid, and 15% Remote job distribution.

Software Development Engineer in Test

Clockwork.io

On-site

CA$130K - CA$175K/yr

Full-time

Posted 20 days ago


Job description

Summary

We are looking for a Software Development Engineer in Test who thrives in fast-paced startup environments and is excited to be a co-owner and major contributor to our testing and CI/CD infrastructure. You will play a critical role in ensuring the quality, performance, and reliability of every release before it reaches our customers.

Our software runs on production GPU and CPU clusters across Ethernet, RDMA/RoCE, and InfiniBand networks. Testing here means standing up real multi-node clusters, running real workloads against them, and measuring what happens. This is a hands-on engineering role, not a manual QA position. You will be deeply embedded in our engineering team, responsible for building robust automation, maintaining pipelines, and contributing directly to the product during quieter times. A meaningful part of this role is DevOps work, not just test authoring.

Responsibilities
  • Design, build, and maintain integration, regression, and end-to-end tests for distributed systems running on Kubernetes, Slurm, and bare-metal GPU/CPU clusters
  • Extend our in-house end-to-end test automation framework and share ownership of the CI/CD testing infrastructure and pipeline
  • Build and run performance and scale benchmarks, and catch regressions before they ship
  • Automate cluster lifecycle and environment management across cloud and bare-metal environments
  • Ensure all code releases meet high quality standards before shipping to customers
  • Collaborate closely with engineers to identify gaps in test coverage and build tools or frameworks to fill them
  • Investigate and diagnose build failures, flaky tests, and regressions
  • Contribute to engineering efforts beyond QA when appropriate (feature development, tooling, etc.)
Qualifications
  • 3+ years of experience in a Test Engineering, DevOps, or QA role with a strong technical background
  • Strong Python; comfortable reading and debugging Go
  • Demonstrated expertise in building automated test scripts and frameworks from scratch
  • Proven experience with CI/CD tools such as GitHub Actions, Jenkins, or GitLab CI, and testing frameworks such as pytest and Playwright
  • Experience with public cloud infrastructure (GCP, AWS, or Azure)
  • Familiarity with release management and deployment processes, including versioning, release candidates, and rollout verification
  • Working knowledge of Linux, containerization, and orchestration (Docker, Kubernetes, Helm)
  • Strong debugging, problem-solving, and communication skills
  • A startup mindset, self-motivated, adaptable, and eager to take ownership
Nice to have
  • Experience with AI/ML infrastructure: GPUs, RDMA/RoCE or InfiniBand, NCCL, DCGM, PyTorch training jobs
  • Workload managers and schedulers (Slurm, Kubeflow/PyTorchJob)
  • Build systems at scale (Bazel), infrastructure-as-code (Terraform)
  • Chaos or fault-injection testing of distributed systems
Why Join Us
  • Be part of a foundational team shaping the future of infrastructure software
  • Work onsite in a collaborative environment in Palo Alto
  • Opportunity to grow into broader engineering roles

Enjoy

  • Challenging projects.
  • A friendly and inclusive workplace culture.
  • Competitive compensation.
  • A great benefits package.
  • Catered lunch.

Compensation for this position will vary based on the skills and experience you bring, as well as internal equity considerations. For candidates hired at the posted level, the expected base salary range is $130,000 - $175,000. The offered compensation package may also include stock options or other equity awards, subject to Clockwork's equity program and applicable approvals.
In addition to cash compensation, this role is eligible to participate in the company's equity program, which may include stock options granted in accordance with the company's equity plan and subject to approval and applicable vesting schedules.