1

Gpu Performance Engineer Jobs in Massachusetts (NOW HIRING)

AI HPC Infrastructure Engineer

Boston, MA ยท Hybrid

$117K - $153K/yr

The AI HPC Infrastructure Engineer owns the operation, performance, and growth of a hybrid high-performance computing (HPC) and AI/GPU infrastructure environment. The engineer maintains the Linux ...

Senior RHEL Systems Engineer

Boston, MA ยท On-site

$113K - $155K/yr

Senior RHEL Systems Engineer Job Location: Boston, MA Job Type: Contract * Installing, configuring ... Experience supporting GPU workstation setups. * A strong understanding of HPC performance testing.

Staff Embedded ML Engineer, Edge AI

Boston, MA ยท On-site

$142K - $187K/yr

Optimize end-to-end inference performance across CPU/DSP/NPU/GPU (as applicable): latency ... Partner closely with ML engineers to translate model changes into deployment impact; provide ...

... performance * Collaborate closely with teams provisioning and operating largeโ€‘scale GPU/TPU ... Lead, mentor, and grow a team of data and infrastructure engineers * Define and track KPIs for data ...

Senior HPC and Quantum Systems Engineer

Westford, MA ยท Hybrid

$108K - $148K/yr

An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can ... This role sits at the intersection of data center infrastructure, high-performance computing, and ...

Analyze model performance and identify improvements. * Research/evaluate emerging technologies to ... PyTorch, * GPU training, * acceleration and inference optimization, * TensorRT/CUDA, * CI/CD ...

Senior Software Engineer, Next Gen Compute

Boston, MA ยท Hybrid

$133K - $175K/yr

Analyze ML workload performance on a variety of hardware processors, optimize ML models, improve ML ... Optimize the utilization of GPU/NPU resources and sharing of GPU/NPU access across multiple ...

Showing results 21-40

Gpu Performance Engineer information

What are some common challenges faced by a GPU performance engineer when optimizing graphics workloads?

GPU Performance Engineers often encounter challenges such as identifying performance bottlenecks within complex graphics pipelines, balancing resource utilization, and achieving optimal frame rates across diverse hardware configurations. They must use specialized profiling tools and collaborate closely with developers, driver engineers, and QA teams to address issues like memory bandwidth limitations or shader inefficiencies. Staying updated with rapidly evolving GPU architectures and optimizing for both current and next-generation hardware are also key aspects of the role.

What is a GPU performance engineer?

A GPU Performance Engineer is a specialist who analyzes, optimizes, and improves the performance of graphics processing units (GPUs). They work on identifying bottlenecks, optimizing code, and ensuring that GPU hardware and software deliver maximum efficiency and speed. Their role may involve working with drivers, firmware, and applications to enhance graphics and compute workloads. This job is essential in industries like gaming, AI, and high-performance computing where GPU efficiency directly impacts user experience and system performance.

What are the key skills and qualifications needed to thrive as a GPU performance engineer, and why are they important?

To thrive as a GPU Performance Engineer, you need a strong background in computer architecture, programming (C/C++), and a degree in computer science, electrical engineering, or a related field. Proficiency with GPU profiling tools (e.g., NVIDIA Nsight, AMD Radeon GPU Profiler), performance analysis frameworks, and parallel computing libraries like CUDA or OpenCL is typically required. Analytical thinking, problem-solving abilities, and effective communication are crucial soft skills for collaborating with developers and debugging performance bottlenecks. These skills and qualities are essential for optimizing GPU performance, ensuring efficient software-hardware interaction, and delivering high-quality graphics or compute solutions.

What is the difference between Gpu Performance Engineer vs Gpu Hardware Engineer?

AspectGpu Performance EngineerGpu Hardware Engineer
Primary FocusOptimizing GPU performance, benchmarking, and tuning softwareDesigning, developing, and testing GPU hardware components
Required SkillsProgramming, performance analysis, GPU architecture knowledgeHardware design, circuit analysis, FPGA/ASIC experience
Work EnvironmentSoftware development teams, labs for testing performanceHardware labs, manufacturing facilities, R&D centers
Common CertificationsNone specific, often requires computer engineering or related degreesElectrical engineering, VLSI design certifications

The Gpu Performance Engineer primarily focuses on optimizing and testing GPU software performance, while the Gpu Hardware Engineer designs and develops the physical GPU components. Both roles require a strong background in computer engineering, but differ in their core responsibilities and work environments.

What job categories do people searching Gpu Performance Engineer jobs in Massachusetts look for? The top searched job categories for Gpu Performance Engineer jobs in Massachusetts are:
What cities in Massachusetts are hiring for Gpu Performance Engineer jobs? Cities in Massachusetts with the most Gpu Performance Engineer job openings:
Infographic showing various Gpu Performance Engineer job openings in Massachusetts as of August 2026, with employment types broken down into 100% Full Time. Highlights an 49% In-person, and 51% Remote job distribution.

AI HPC Infrastructure Engineer

Analysis Group, Inc.

Boston, MA โ€ข Hybrid

$117K - $153K/yr

Full-time

Posted 11 days ago


Job description

Overview

Analysis Group is one of the largest international economics consulting firms, with more than 1,500 professionals across 15 offices in North America, Europe, and Asia. Since 1981, we have provided expertise in economics, finance, health care analytics, and strategy to top law firms, Fortune Global 500 companies, and government agencies worldwide. Our internal experts, together with our network of affiliated experts from academia, industry, and government, offer our clients exceptional breadth and depth of expertise.

The AI HPC Infrastructure Engineer owns the operation, performance, and growth of a hybrid high-performance computing (HPC) and AI/GPU infrastructure environment. The engineer maintains the Linux-based clustered computing platform that supports both traditional HPC/analytical workloads and large-scale AI/ML training and inference, ensuring systems run efficiently, GPUs and other accelerators are current and well-utilized, and operations are monitored, documented, and reported - including change management and performance statistics - across both domains.

Essential Job Functions and Responsibilities

  • Maintain, tune, and manage the analytical and AI computing environment for researchers and data scientists, including Posit Workbench (RStudio Server Pro) environments.
  • Optimize systems and infrastructure performance using parallelization technologies (MPI, OpenMP) and distributed/multi-GPU training strategies (e.g., PyTorch Distributed, Horovod, DeepSpeed).
  • Design, deploy, and maintain GPU-accelerated compute infrastructure for large-scale model training and inference.
  • Manage GPU scheduling, multi-tenancy, and utilization across SLURM and/or Kubernetes-based environments.
  • Administer the NVIDIA software stack - drivers, CUDA, cuDNN, NCCL - and coordinate firmware and health monitoring across GPU fleets.
  • Tune and optimize LLM training and inference performance - including batching, quantization, KV-cache utilization, parallelism strategies, and throughput/latency across GPU clusters.
  • Build and maintain MLOps pipelines for model training, versioning, deployment, and monitoring (e.g., MLflow, Kubeflow).
  • Manage container orchestration and runtimes (Docker, Kubernetes, Singularity/Apptainer) supporting both HPC jobs and ML workloads.
  • Manage access authentication including PAM, LDAP integration, and single sign-on.
  • Design and develop scripts for system administration, automating tasks, monitoring, and usage reporting across HPC and AI resources.
  • Manage high-performance storage and data pipelines for AI training datasets and HPC workloads, primarily on GPFS (IBM Spectrum Scale).
  • Troubleshoot, isolate, and resolve application, systems, and other technical problems (hardware, software, network, and GPU-specific issues).
  • Develop and implement backup and recovery programs.
  • Research, deploy, and manage general infrastructure, including development of policies and procedures for both HPC and AI/ML environments.
  • Migrate data from heterogeneous environments to Linux, on-prem clusters, or cloud.
  • Collaborate with data scientists and ML engineers to support the model development lifecycle and translate research needs into infrastructure requirements.
  • Evaluate emerging AI hardware, accelerators, and cloud AI services, and recommend adoption where beneficial.
  • Monitor performance, troubleshoot problem areas, and provide statistics and reports across compute, storage, and network.
  • Create and maintain documentation related to system configuration, processes, change management, inventory, and service records.
  • Ensure continuous network connectivity of all equipment.
  • Conduct research and report on products, services, protocols, and standards to remain abreast of developments in HPC and AI infrastructure.
  • Participate in a 24x7 on-call rotation; troubleshoot and resolve issues remotely or onsite as necessary.

Qualifications

  • Bachelor's degree required; degree in computer science, electrical engineering, or a related field preferred.
  • A minimum of 5 years of experience as a hands-on Linux Systems Administrator in a research, HPC, or production setting.
  • An ideal candidate will have 5 to 10 years of substantive relevant experience.ย 
  • Experience managing Posit Workbench (RStudio Server Pro), Python, and R environments; strong Posit Workbench administration experience is a significant plus.
  • Experience with SLURM, Platform LSF, or other job schedulers required; experience scheduling GPU resources strongly preferred.
  • Hands-on experience with NVIDIA GPU infrastructure and software stack (CUDA, cuDNN, NCCL, NVIDIA GPU Operator) strongly preferred.
  • Experience with Kubernetes and container orchestration for AI/ML workloads highly desired.
  • Familiarity with ML/AI frameworks (PyTorch, TensorFlow) and distributed training patterns highly desired.
  • Experience with MLOps tooling (MLflow, Kubeflow, Weights & Biases, or similar) is a plus.
  • Experience with Bright Cluster Manager is highly desired.
  • Experience with Ansible is highly desired.
  • Experience with containerization (Docker, Singularity/Apptainer) is highly desired.
  • Proficiency with remote access technologies and tools such as RDP, SSH, and emulation software
  • Hands-on experience with GPFS (IBM Spectrum Scale) required.
  • Demonstrated experience tuning LLM training and/or inference performance (e.g., batching, quantization, KV-cache management, parallelism strategies) required.
  • Experience with AI Gateways (e.g., LiteLLM, Kong AI Gateway, Portkey, or similar) is a very nice to have.
  • Excellent hardware troubleshooting experience, including GPU-specific diagnostics.
  • Knowledge of applicable data privacy practices and laws.
  • Strong interpersonal, written, and oral communication skills.
  • Highly self-motivated and directed, with keen attention to detail.
  • Proven analytical and problem-solving abilities.
  • Strong customer service orientation.
  • Experience working in a collaborative environment.
  • An inclusive and growth-oriented mindset, strong interpersonal skills, and an ability to work across functions.
  • To the extent permitted by applicable law, eligible candidates must be authorized to work in the United States, without sponsorship or restriction, now and in the future.

Analysis Group embraces equal opportunity. We are committed to building teams that bring a variety of backgrounds, perspectives, and skills, as we believe that a strong and inclusive workforce directly supports our goal of providing the highest-quality work. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, or any other class protected under applicable federal, state, or local law, and we encourage candidates of all backgrounds to apply.

Analysis Group offers competitive compensation and a comprehensive benefits package. The estimated salary range for this position is $150,000-$170,000. Compensation offered will be based on a number of factors including work experience, education, and skill level. This role is eligible for a discretionary annual bonus that is determined in large part by individual performance. To learn more about our benefit offerings, clickย here.

#LI-Hybrid

Privacy Notice

For information about Analysis Group's privacy practices, please refer to the applicable Analysis Groupย privacy policy.

  • Equal Opportunity Employer/Protected Veterans/Individuals with Disabilities.
  • Please view the EEOC's "Know Your Rights" poster here.
Employment Type: OTHER