1

Gpu Performance Engineer Jobs in Montana (NOW HIRING)

We're a team of engineers and scientists with deep backgrounds in ML infrastructure & research ... Understanding of cluster scheduling, networking bottlenecks, and GPU/TPU performance optimization ...

Gpu Performance Engineer information

What is a GPU performance engineer?

A GPU Performance Engineer is a specialist who analyzes, optimizes, and improves the performance of graphics processing units (GPUs). They work on identifying bottlenecks, optimizing code, and ensuring that GPU hardware and software deliver maximum efficiency and speed. Their role may involve working with drivers, firmware, and applications to enhance graphics and compute workloads. This job is essential in industries like gaming, AI, and high-performance computing where GPU efficiency directly impacts user experience and system performance.

What are some common challenges faced by a GPU performance engineer when optimizing graphics workloads?

GPU Performance Engineers often encounter challenges such as identifying performance bottlenecks within complex graphics pipelines, balancing resource utilization, and achieving optimal frame rates across diverse hardware configurations. They must use specialized profiling tools and collaborate closely with developers, driver engineers, and QA teams to address issues like memory bandwidth limitations or shader inefficiencies. Staying updated with rapidly evolving GPU architectures and optimizing for both current and next-generation hardware are also key aspects of the role.

What are the key skills and qualifications needed to thrive as a GPU performance engineer, and why are they important?

To thrive as a GPU Performance Engineer, you need a strong background in computer architecture, programming (C/C++), and a degree in computer science, electrical engineering, or a related field. Proficiency with GPU profiling tools (e.g., NVIDIA Nsight, AMD Radeon GPU Profiler), performance analysis frameworks, and parallel computing libraries like CUDA or OpenCL is typically required. Analytical thinking, problem-solving abilities, and effective communication are crucial soft skills for collaborating with developers and debugging performance bottlenecks. These skills and qualities are essential for optimizing GPU performance, ensuring efficient software-hardware interaction, and delivering high-quality graphics or compute solutions.

What is the difference between Gpu Performance Engineer vs Gpu Hardware Engineer?

AspectGpu Performance EngineerGpu Hardware Engineer
Primary FocusOptimizing GPU performance, benchmarking, and tuning softwareDesigning, developing, and testing GPU hardware components
Required SkillsProgramming, performance analysis, GPU architecture knowledgeHardware design, circuit analysis, FPGA/ASIC experience
Work EnvironmentSoftware development teams, labs for testing performanceHardware labs, manufacturing facilities, R&D centers
Common CertificationsNone specific, often requires computer engineering or related degreesElectrical engineering, VLSI design certifications

The Gpu Performance Engineer primarily focuses on optimizing and testing GPU software performance, while the Gpu Hardware Engineer designs and develops the physical GPU components. Both roles require a strong background in computer engineering, but differ in their core responsibilities and work environments.

What are popular job titles related to Gpu Performance Engineer jobs in Montana?

For Gpu Performance Engineer jobs in Montana, the most frequently searched job titles are:

What job categories do people searching Gpu Performance Engineer jobs in Montana look for?

The top searched job categories for Gpu Performance Engineer jobs in Montana are:

Member of Technical Staff - Foundations

Tzafon

Zurich, MT • On-site

$200K - $500K/yr

Full-time

Medical, Dental, Vision, Retirement

Re-posted 13 days ago


Job description

Tzafon is a foundation model lab building scalable compute systems and advancing machine intelligence, with offices in San Francisco, Zurich & Tel Aviv. We've raised over $12m in funding to advance our mission of expanding the frontiers of machine intelligence.
We're a team of engineers and scientists with deep backgrounds in ML infrastructure & research. Founded by IOI and IMO medalists, PhDs, and alumni from leading tech companies, such as Google Deepmind, Character, and NVIDIA, we train models and build infrastructure for swarms of agents to automate work across real-world environments.
You'll work between our product and post-training teams to ship Large Action Models that actually work. Build evals, benchmarks, and fine-tuning pipelines. Define what good model behavior means and make it happen at scale.
What you'll do
  • Design and execute large scale training runs on our clusters
  • Build and optimize distributed training infrastructure across massive multi-node systems
  • Implement post-training pipelines at scale
  • Develop data pipelines that process and filter trillions of tokens for pre-training
  • Research and implement architectural improvements, scaling laws, and training optimizations
  • Debug training instabilities, loss spikes, and convergence issues in long-running jobs
  • Build tooling for cluster utilization, fault tolerance, and checkpoint management
  • Write custom CUDA/Triton kernels to optimize critical training operations (attention, normalization, activations)
  • Collaborate on research that advances the state of the art in foundation model training

We're looking for
  • Deep experience pre-training or post-training foundation models on large clusters
  • Expert-level at Python and ML frameworks (PyTorch, JAX, Torchtitan)
  • Strong systems skills: distributed training, FSDP/ZeRO, tensor parallelism, pipeline parallelism
  • Experience writing performant CUDA or Triton kernels for ML workloads
  • Track record of running stable multi-week training jobs and debugging distributed training failures
  • Understanding of cluster scheduling, networking bottlenecks, and GPU/TPU performance optimization

Preferred Experience
  • Trained foundation models at major AI labs (OpenAI, Anthropic, Google DeepMind, Meta, xAI, etc.)
  • Worked on large scale RL runs
  • Optimized critical training kernels (FlashAttention, fused optimizers, custom kernels)
  • Published research at top ML conferences (NeurIPS, ICML, ICLR)
  • Contributions to open source ML infrastructure (PyTorch, JAX, vLLM, etc.)
  • Experience with training data pipelines, data quality research, or synthetic data generation

Life at Tzafon
  • Full medical, dental, and vision coverage, plus 401(k) in the us
  • Office in SF, Zurich, and Tel Aviv
  • Early-stage equity in a future-defining company

Visa sponsorship: We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.
Compensation starts at $200k-$500k + equity package, depending on experience & location.
We also offer a referral bonus of $5k for referral of successful hires (send to careers@tzafon.ai).