1

Deep Learning Performance Architect Jobs (NOW HIRING)

Performance Architect, AI HW

$170K/yr

... deep learning workloads into architectural insight and measurable design tradeoffs. • Curious ... in high-performance AI systems. • Benchmark and analyze complex AI workloads across single and ...

Lead Performance Architect

Orlando, FL · On-site

$94K - $141K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Performance Architects focus on performance outcomes-not just training-by evaluating root causes ... Collaborate with Learning Design and Delivery teams to develop and deploy performance solutions.

We are now looking for a Senior Performance Architect for Nemotron! At NVIDIA, we are redefining ... Experience with deep learning frameworks like PyTorch, TRT-LLM, VLLM, SGLang * A Growth mindset and ...

Lead Performance Architect

Orlando, FL · On-site

$94K - $141K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Performance Architects focus on performance outcomes-not just training-by evaluating root causes ... Collaborate with Learning Design and Delivery teams to develop and deploy performance solutions.

We are now looking for a Senior Performance Architect for Nemotron! At NVIDIA, we are redefining ... Experience with deep learning frameworks like PyTorch, TRT-LLM, VLLM, SGLang * A Growth mindset and ...

We are now looking for a Senior Performance Architect for Nemotron! At NVIDIA, we are redefining ... Experience with deep learning frameworks like PyTorch, TRT-LLM, VLLM, SGLang * A Growth mindset and ...

The NVIDIA GPU Architecture group is looking for world class architects and software developers to ... highest performance in the world for deep learning and parallel processing algorithms. We are ...

The NVIDIA GPU Architecture group is looking for world class architects and software developers to ... highest performance in the world for deep learning and parallel processing algorithms. We are ...

Senior CPU Performance Architect

Hillsboro, OR · On-site

$181K/yr

... deep learning (DL), high-performance computing (HPC), cloud service providers (CSP), gaming ... Come join the CPU performance architecture team and help us push performance boundaries for all our ...

Showing results 21-40

Deep Learning Performance Architect information

See salary details

$156.5K

$168K

How much do deep learning performance architect jobs pay per year?

As of Aug 13, 2026, the average yearly pay for deep learning performance architect in the United States is $167,842.00, according to ZipRecruiter salary data. Most workers in this role earn between $167,000.00 and $167,000.00 per year, depending on experience, location, and employer.

What are the key skills and qualifications needed to thrive as a deep learning performance architect?

To thrive as a Deep Learning Performance Architect, you need a strong background in computer science, deep learning frameworks, parallel computing, and optimization techniques, typically supported by a relevant degree and experience in AI or high-performance computing. Familiarity with tools such as TensorFlow, PyTorch, CUDA, and profiling or benchmarking systems is essential. Analytical problem-solving, effective communication, and a collaborative mindset help professionals excel in cross-functional teams and resolve complex performance bottlenecks. These skills are vital for optimizing AI workloads, ensuring scalability, and maximizing the efficiency of deep learning models in production environments.

What is a deep learning performance architect?

A Deep Learning Performance Architect is a specialized professional who designs, analyzes, and optimizes the performance of deep learning systems and models. They work to improve the efficiency, speed, and scalability of machine learning algorithms on various hardware platforms such as GPUs, TPUs, and CPUs. Their role often involves collaborating with software engineers and data scientists to identify bottlenecks and implement solutions that enhance computational capabilities for AI workloads. By doing so, they ensure that deep learning applications run faster and more efficiently, making the best use of available resources.

What is the difference between Deep Learning Performance Architect vs Machine Learning Engineer?

AspectDeep Learning Performance ArchitectMachine Learning Engineer
CredentialsAdvanced degrees in AI, deep learning, or related fields; certifications in deep learning frameworksDegrees in computer science, data science, or related fields; certifications in machine learning tools
Work EnvironmentResearch labs, AI development teams, performance optimization settingsData-driven projects, model development, deployment environments
Industry UsageTech companies, AI research firms, organizations focusing on deep learning optimizationTech companies, startups, enterprises applying machine learning solutions

The Deep Learning Performance Architect specializes in optimizing deep learning models for efficiency and scalability, focusing on hardware and software performance. In contrast, Machine Learning Engineers develop, train, and deploy machine learning models across various applications. While both roles require strong technical skills, the Architect emphasizes performance tuning and system optimization, whereas the Engineer focuses on model development and implementation.

What are some common challenges faced by deep learning performance architects when optimizing large-scale neural network models?

Deep Learning Performance Architects often encounter challenges such as balancing model accuracy with computational efficiency, managing memory constraints on specialized hardware, and optimizing inference or training speed across different platforms. They frequently need to profile and analyze bottlenecks at both the algorithmic and hardware levels, often requiring close collaboration with software engineers and hardware designers. Staying current with rapidly evolving deep learning frameworks and hardware accelerators is also essential to ensure optimal performance and scalability.
More about Deep Learning Performance Architect jobs
What job categories do people searching Deep Learning Performance Architect jobs look for? The top searched job categories for Deep Learning Performance Architect jobs are:
Infographic showing various Deep Learning Performance Architect job openings in the United States as of August 2026, with employment types broken down into 1% As Needed, 87% Full Time, 10% Part Time, and 2% Contract. Highlights an 91% Physical, 2% Hybrid, and 7% Remote job distribution, with an average salary of $167,842 per year, or $80.7 per hour.

$170K/yr

Full-time

Re-posted 5 days ago


Job description

Job Summary:
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. As an AI Performance Architect, you will model, analyze, and optimize AI workloads on the Tensix architecture, ensuring design decisions deliver measurable performance gains.
Responsibilities:
• Benchmark and analyze complex AI workloads across single and multi-node hardware configurations to guide next-gen architecture.
• Develop and maintain performance models, simulators, and micro-benchmark suites to drive feature evaluation and design optimization.
• Conduct detailed PPA (Performance, Power, Area) studies to assess design tradeoffs and inform hardware-software co-design decisions.
• Collaborate closely with RTL, Compiler, and Runtime teams to instrument and correlate performance models with silicon results.
Qualifications:
Required:
• Deeply analytical engineer with strong intuition for AI workload behavior and system-level performance bottlenecks.
• Experienced in C++ and Python for simulation, modeling, and performance analysis across heterogeneous compute systems.
• Adept at bridging software and hardware teams—translating deep learning workloads into architectural insight and measurable design tradeoffs.
• Curious, data-driven, and comfortable pushing the limits of efficiency, scalability, and accuracy in high-performance AI systems.
• Benchmark and analyze complex AI workloads across single and multi-node hardware configurations to guide next-gen architecture.
• Develop and maintain performance models, simulators, and micro-benchmark suites to drive feature evaluation and design optimization.
• Conduct detailed PPA (Performance, Power, Area) studies to assess design tradeoffs and inform hardware-software co-design decisions.
• Collaborate closely with RTL, Compiler, and Runtime teams to instrument and correlate performance models with silicon results.
Company:
Tenstorrent develops AI hardware and software solutions for data processing and machine learning application. Founded in 2016, the company is headquartered in Toronto, CAN, with a team of 501-1000 employees. The company is currently Late Stage.