1

Internship Forward Deployed Software Engineer Jobs in Toronto, ON

You'll be part of a collaborative, forward-thinking team rewriting a legacy, monolithic application ... Deploy, manage and scale applications in cloud environments using Kubernetes. * Execute technical ...

Design, develop, write comprehensive automated tests for, and deploy robust software applications ... Our forward-looking companies lead the way in software-powered workflow solutions, data-driven ...

... fully deployed, production-ready application. You have done time in both larger engineering ... Must-Haves * 2+ years of professional software engineering experience across large organizations or ...

... fully deployed, production-ready application. You have done time in both larger engineering ... Must-Haves * 2+ years of professional software engineering experience across large organizations or ...

We are looking for an enthusiastic and motivated software engineer to join our marketplace teams ... internships, personal projects, or academic projects are highly valued). * Technical Skills ...

Principal Software Engineer

Toronto, ON · On-site

CA$220K - CA$300K/yr

... deployed today, running in production at some of the largest carriers in North America. We are ... Join us at Owl.co and become an integral part of a forward-thinking team dedicated to transforming ...

Software Development * Develop clean, maintainable, and well documented code using .NET, C#, Java ... Develop and deploy solutions in Azure using Azure DevOps (TFS) pipelines. * Use Terraform (or other ...

We're looking for an Software Engineer to join our Automotive Finance Engineering team. You will ... Develop and deploy solutions in Azure using Azure DevOps (TFS) pipelines. * Use Terraform (or other ...

Engineer, Software

Toronto, ON

CA$185K - CA$225K/yr

Every day, we work to reinvent and lead our industry forward by thinking bigger and challenging the ... Improve developer velocity - own the build, deploy, and testing pipeline and continuously reduce ...

Showing results 21-40

Internship Forward Deployed Software Engineer information

What is the difference between Internship Forward Deployed Software Engineer vs Software Engineer Intern?

AspectInternship Forward Deployed Software EngineerSoftware Engineer Intern
CredentialsTypically pursuing or holding a bachelor's or master's in CS or related fieldUsually students or recent graduates in CS or related fields
Work EnvironmentHands-on, real-world deployment, close client interaction, fast-pacedLearning-focused, project-based, mentorship-driven
Employer & Industry UsageTech companies, especially in cloud, AI, or enterprise solutionsTech companies, startups, or corporate R&D teams

The Internship Forward Deployed Software Engineer role involves working directly on deployed systems with real clients, often requiring more advanced skills and understanding of deployment processes. In contrast, a Software Engineer Intern typically focuses on learning, supporting ongoing projects, and gaining industry experience. Both roles are valuable entry points but differ mainly in responsibility level and work scope.

Infographic showing various Internship Forward Deployed Software Engineer job openings in Toronto, ON as of August 2026, with employment types broken down into 100% Full Time. Highlights an 80% In-person, and 20% Remote job distribution.

Senior Software Engineer, AI Inference

Nvidia

Toronto, ON • Hybrid

Full-time

Re-posted 7 days ago


Nvidia rating

9.6

Company rating: 9.6 out of 10

Based on 17 frontline employees who took The Breakroom Quiz

7th of 244 rated software companies


Job description

Help us push the boundaries of AI inference at NVIDIA - where your systems expertise shapes both the technology and the teams building on top of it!

We're looking for a Senior Software Engineer to work at the frontier of large-scale LLM serving, partnering directly with some of the world's most technically demanding customers to unlock the full performance potential of NVIDIA's inference stack. In this role, you'll combine deep systems knowledge with hands-on customer engagement - profiling real deployments, benchmarking across GPU clusters, and turning insights into improvements that ripple across the open-source ecosystem. Do you love digging into performance problems that don't have obvious answers, and want your work to have an impact far beyond a single codebase? We'd love to talk. Unlike traditional customer-facing engineering roles, we expect you to go far deeper - contributing to vLLM, NVIDIA Dynamo, and the tooling that makes every engineer on your team more effective.

What You'll be doing:

  • Work directly with customer engineering teams through long-term technical partnerships, understanding their LLM serving architectures and performance goals, then designing and implementing end-to-end benchmarking campaigns across Kubernetes and Slurm environments to surface actionable insights.

  • Set up and operate vLLM serving deployments on GPU clusters, tuning configurations for throughput, latency, and efficiency - and collect Nsight Systems / Nsight Compute profiling traces to identify performance gaps relative to reference frameworks.

  • Develop detailed performance plans based on profiling findings and collaborate with NVIDIA's kernel engineering and OSS vLLM teams to drive improvements that benefit both your customers and the broader community.

  • Build internal tools, benchmarking harnesses, and automation pipelines that raise the productivity of your teammates and customers alike - with a multiplier attitude that makes everyone around you more effective.

  • Document architectures, findings, and recommendations with clarity for technical audiences, and contribute improvements back to vLLM and related open-source projects where appropriate.

What We Need to See:

  • Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, or equivalent experience.

  • 5+ years of industry experience building and operating complex, production-grade software systems, with strong instincts for how systems behave at scale.

  • Hands-on experience deploying and operating LLM inference workloads - particularly with vLLM - including configuration, optimization, and debugging in real-world environments.

  • Proficiency with container orchestration (Kubernetes) and HPC scheduling (Slurm) for running GPU-accelerated workloads.

  • Solid understanding of LLM serving fundamentals: batching strategies (continuous batching, chunked prefill), KV cache management, and tensor/pipeline parallelism.

  • Familiarity with GPU performance analysis: memory hierarchy, utilization, roofline modeling, and profiling with Nsight Systems or Nsight Compute.

  • Strong written and verbal communication skills, with the ability to present technical findings clearly to both engineering teams and leadership - and to navigate ambiguous, open-ended customer problems.

Ways to Stand Out from the Crowd:

  • Experience with NVIDIA Dynamo or other disaggregated inference serving frameworks.

  • Contributions to open-source inference or ML systems projects, particularly vLLM or SGLang - please include links to relevant pull requests or artifacts.

  • Background with ML compilers or GPU kernel development (Triton, CUTLASS, TorchInductor).

  • Experience building developer tools or internal platforms that meaningfully improved team productivity.

  • Prior experience in a customer-facing or forward-deployed engineering capacity within a technical product organization.

Widely considered to be one of the technology world's most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family www.nvidiabenefits.com/

#LI-Hybrid

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 135,000 CAD - 185,000 CAD for Level 3, and 170,000 CAD - 220,000 CAD for Level 4.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until April 14, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.


What Nvidia employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom


Nvidia logo

About Nvidia

Sourced by ZipRecruiter

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology--and amazing people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what's never been done before takes vision, innovation, and the world's best talent.

Industry

Computer and electronic product manufacturing

Company size

10,000+ Employees

Headquarters location

Santa Clara, CA, US

Year founded

1993