1

Gpu Performance Engineer Jobs in Phoenix, AZ (NOW HIRING)

AI Performance Engineer

Phoenix, AZ · On-site

$55 - $60/hr

AI Performance Engineer Client: Infosys/Microsoft Pheonix, AZ// Seattle, WA Rate: $55-60/hr C2C AI ... of GPU utilization, memory bandwidth, and data pipelines. Responsibilities Benchmark AI models ...

You will apply your expertise in CUDA, C++, and GPU programming to analyze, optimize, and improve high-performance GPU kernels. No prior AI experience is required. Key Responsibilities * Analyze ...

You will apply your expertise in CUDA, C++, and GPU programming to analyze, optimize, and improve high-performance GPU kernels. No prior AI experience is required. Key Responsibilities * Analyze ...

CUDA Developer - Remote

Phoenix, AZ · Remote

$60 - $100/hr

You will apply your expertise in CUDA, C++, and GPU programming to analyze, optimize, and improve high-performance GPU kernels. No prior AI experience is required. Key Responsibilities * Analyze ...

You will apply your expertise in GPU programming, performance optimization, and C++ development to build high-performance solutions. Key Responsibilities * Design, implement, and optimize GPU ...

You will apply your expertise in GPU programming, performance optimization, and C++ development to build high-performance solutions. Key Responsibilities * Design, implement, and optimize GPU ...

Machine Learning Engineer

Phoenix, AZ · On-site

$130K - $150K/yr

Sr. Machine Learning Engineer Salary Range: $130k to $150k Our client is seeking a Sr. Machine ... Profile and optimize system performance across CPU, GPU, memory, and I/O subsystems. GPU ...

Senior Software Engineer

Chandler, AZ · On-site

$150 - $200/hr

... performance of existing functionality * Decomposes functional requirements into well-defined tasks ... Knowledge of GPU architectures and Graphics pipeline * Serving as a lead software engineer for a ...

New

Senior Software Engineer

Tempe, AZ · On-site

$117K - $154K/yr

You will build high-performance, scalable, and maintainable software components that serve as the ... Collaborate with AI, vision, and embedded teams to integrate and optimize GPU-accelerated ...

Hardware Systems Engineer

Phoenix, AZ · On-site

$122K - $161K/yr

Work closely with firmware development teams, to drive device performance via understanding of the ... CPU or GPU microarchitecture, batteries, power management, analog design, audio, wireless ...

Data Center Technician

Phoenix, AZ · On-site

$30 - $36/hr

Our network of 1,000+ field engineers operates globally, tackling the most complex deployments in ... performance and business demand. Key Responsibilities GPU Infrastructure & Hardware Management ...

Data Center Technician

Phoenix, AZ · On-site

$30 - $36/hr

Our network of 1,000+ field engineers operates globally, tackling the most complex deployments in ... performance and business demand. Key Responsibilities GPU Infrastructure & Hardware Management ...

Avionics Software Engineer

Phoenix, AZ · Hybrid

$100K - $185K/yr

... performance and avoid runtime stalls. · Perform software integration tests, unit tests, and other ... Preferred Qualifications & Skills: · Display integration experience (touchscreen, HUD, and/or GPU ...

... performance and avoid runtime stalls. · Perform software integration tests, unit tests, and other ... Preferred Qualifications & Skills: · Display integration experience (touchscreen, HUD, and/or GPU ...

next page

Showing results 1-20

Gpu Performance Engineer information

See Phoenix, AZ salary details

$10

$59

$97

How much do gpu performance engineer jobs pay per hour?

As of Sep 7, 2026, the average hourly pay for gpu performance engineer in Phoenix, AZ is $59.68, according to ZipRecruiter salary data. Most workers in this role earn between $48.94 and $67.55 per hour, depending on experience, location, and employer.

What is a GPU performance engineer?

A GPU Performance Engineer is a specialist who analyzes, optimizes, and improves the performance of graphics processing units (GPUs). They work on identifying bottlenecks, optimizing code, and ensuring that GPU hardware and software deliver maximum efficiency and speed. Their role may involve working with drivers, firmware, and applications to enhance graphics and compute workloads. This job is essential in industries like gaming, AI, and high-performance computing where GPU efficiency directly impacts user experience and system performance.

What are some common challenges faced by a GPU performance engineer when optimizing graphics workloads?

GPU Performance Engineers often encounter challenges such as identifying performance bottlenecks within complex graphics pipelines, balancing resource utilization, and achieving optimal frame rates across diverse hardware configurations. They must use specialized profiling tools and collaborate closely with developers, driver engineers, and QA teams to address issues like memory bandwidth limitations or shader inefficiencies. Staying updated with rapidly evolving GPU architectures and optimizing for both current and next-generation hardware are also key aspects of the role.

What are the key skills and qualifications needed to thrive as a GPU performance engineer, and why are they important?

To thrive as a GPU Performance Engineer, you need a strong background in computer architecture, programming (C/C++), and a degree in computer science, electrical engineering, or a related field. Proficiency with GPU profiling tools (e.g., NVIDIA Nsight, AMD Radeon GPU Profiler), performance analysis frameworks, and parallel computing libraries like CUDA or OpenCL is typically required. Analytical thinking, problem-solving abilities, and effective communication are crucial soft skills for collaborating with developers and debugging performance bottlenecks. These skills and qualities are essential for optimizing GPU performance, ensuring efficient software-hardware interaction, and delivering high-quality graphics or compute solutions.

What is the difference between Gpu Performance Engineer vs Gpu Hardware Engineer?

AspectGpu Performance EngineerGpu Hardware Engineer
Primary FocusOptimizing GPU performance, benchmarking, and tuning softwareDesigning, developing, and testing GPU hardware components
Required SkillsProgramming, performance analysis, GPU architecture knowledgeHardware design, circuit analysis, FPGA/ASIC experience
Work EnvironmentSoftware development teams, labs for testing performanceHardware labs, manufacturing facilities, R&D centers
Common CertificationsNone specific, often requires computer engineering or related degreesElectrical engineering, VLSI design certifications

The Gpu Performance Engineer primarily focuses on optimizing and testing GPU software performance, while the Gpu Hardware Engineer designs and develops the physical GPU components. Both roles require a strong background in computer engineering, but differ in their core responsibilities and work environments.

What are popular job titles related to Gpu Performance Engineer jobs in Phoenix, AZ?

For Gpu Performance Engineer jobs in Phoenix, AZ, the most frequently searched job titles are:

What job categories do people searching Gpu Performance Engineer jobs in Phoenix, AZ look for?

The top searched job categories for Gpu Performance Engineer jobs in Phoenix, AZ are:

What cities near Phoenix, AZ are hiring for Gpu Performance Engineer jobs?

Cities near Phoenix, AZ with the most Gpu Performance Engineer job openings:

AI Performance Engineer

Info Way Solutions

Phoenix, AZ • On-site

$55 - $60/hr

Other

Posted 4 days ago


Job description

AI Performance Engineer

Client: Infosys/Microsoft

Pheonix, AZ// Seattle, WA

Rate: $55-60/hr C2C

AI Model tuning:
AI Performance Engineer — Model Optimization & Systems
About the role You will own the question: "How well do our AI models run on our hardware, and how do we make them run better?" You'll benchmark and profile training and inference workloads, identify compute, memory, and I/O bottlenecks, and recommend (and implement) optimizations — from model-level techniques like quantization and batching to system-level tuning of GPU utilization, memory bandwidth, and data pipelines.
Responsibilities
Benchmark AI models (LLMs, vision, multimodal) across hardware configurations; measure latency, throughput, utilization, memory behavior, and scaling efficiency
Profile workloads end-to-end using tools such as Nsight Systems/Compute, PyTorch Profiler, and system telemetry (nvidia-smi, DCGM) to isolate bottlenecks
Build roofline/performance models to quantify achieved vs. theoretical performance and prioritize the highest-impact optimizations
Apply and evaluate optimizations: quantization, pruning, distillation, operator/kernel fusion, graph compilation, KV-cache management, batching strategies, speculative decoding
Recommend hardware/system configurations (GPU selection, memory sizing, interconnect, storage/network I/O) for given model workloads
Establish performance baselines, SLAs, and regression testing so models stay fast as they evolve
Write clear analyses and recommendations for engineering and leadership audiences
Required qualifications
BS/MS in CS, Computer Engineering, EE, or equivalent practical experience
Strong Python; working proficiency in at least one systems language (C++/Rust/C)
Hands-on experience with a deep-learning framework (PyTorch preferred), including model execution, export, and profiling
Demonstrated experience delivering measurable performance improvements in DL training or inference
Solid grounding in computer architecture: memory hierarchy, bandwidth vs. compute limits, parallelism
Ability to reason quantitatively about latency, throughput, batching, memory footprint, and utilization under real workloads
Fluency with Linux and GPU computing environments
Preferred qualifications
GPU programming (CUDA, Triton, ROCm/HIP) and low-level libraries (cuBLAS, cuDNN, CUTLASS)
Inference runtimes/serving engines: TensorRT(-LLM), ONNX Runtime, vLLM, SGLang, Triton Inference Server
LLM inference mechanics: attention, KV caching, prefill vs. decode, continuous batching, speculative decoding
Distributed training/inference: data/tensor/pipeline parallelism, NCCL, InfiniBand/RoCE
Model compression research or MLPerf-style benchmarking experience
Edge/on-device deployment (Jetson, NPUs, Core ML) if your systems include e