1

Pytorch Developer Jobs in Santa Cruz, CA (NOW HIRING)

This can involve anything from digging through PyTorch and machine learning models to determining how to map operations on to our underlying hardware. Responsibilities * Lead compiler engineering ...

Senior GenAI Engineer (Python)

Palo Alto, CA · On-site

$142K - $192K/yr

As a Senior GenAI Developer, you will lead the design, development, and deployment of scalable ... Advanced proficiency in Python and libraries such as PyTorch, Pandas, NumPy, or similar.

Senior GenAI Engineer (Python)

Palo Alto, CA · Hybrid

$142K - $192K/yr

As a Senior GenAI Developer, you will lead the design, development, and deployment of scalable ... Advanced proficiency in Python and libraries such as PyTorch, Pandas, NumPy, or similar.

Staff Engineer, Compiler

San Jose, CA · On-site

$163 - $253/hr

Experience building PyTorch backends for non‑CUDA accelerators (XPU, ROCm, MPS, TPU, custom ... Background in HPC, distributed systems, or NUMA‑aware programming -- anything that built ...

Advanced proficiency in Python and its core data science/ML libraries (e.g., PyTorch, scikit-learn ... Demonstrable, hands-on experience in prompt engineering and/or fine-tuning Large Language Models (e ...

We are hiring a Principal Machine Learning Engineer to serve as the technical lead for our GenAI ... Deep experience with inference frameworks and tools such as PyTorch, CUDA, Triton, TensorRT, Nvidia ...

Its robots are engineered to perform a variety of tasks in the home and commercial markets. Figure ... Proficiency with profiling tools (Nsight Systems/Compute, PyTorch Profiler, HTA, or similar) and ...

Its robots are engineered to perform a variety of tasks in the home and commercial markets. Figure ... Proficiency with profiling tools (Nsight Systems/Compute, PyTorch Profiler, HTA, or similar) and ...

Showing results 21-40

Pytorch Developer information

What is a PyTorch developer?

A PyTorch Developer is a software engineer or data scientist who specializes in using PyTorch, an open-source machine learning library, to build and deploy deep learning models. Their responsibilities typically include designing neural network architectures, training and evaluating models, and optimizing code for performance. PyTorch Developers work in fields such as artificial intelligence, computer vision, and natural language processing, collaborating with teams to solve complex problems using machine learning. They are proficient in Python and have a strong understanding of deep learning concepts. Additionally, they often contribute to research, development, and the deployment of AI solutions in production environments.

What are some common challenges PyTorch developers face when deploying machine learning models to production environments?

Pytorch Developers often encounter challenges when transitioning models from research to production, such as optimizing model performance for inference speed and memory usage, ensuring compatibility with deployment frameworks like TorchScript or ONNX, and managing dependencies across different systems. Additionally, integrating PyTorch models into existing software stacks and maintaining reproducibility can be complex. Collaborating closely with DevOps and data engineering teams is crucial to address these issues and ensure smooth deployment.

What are the key skills and qualifications needed to thrive as a PyTorch developer, and why are they important?

To thrive as a Pytorch Developer, you need strong programming skills in Python, a solid grasp of machine learning concepts, and experience with deep learning frameworks—especially PyTorch itself. Familiarity with tools like CUDA, Jupyter Notebooks, and version control systems (e.g., Git) is typically expected, along with knowledge of cloud platforms or relevant certifications. Problem-solving ability, effective collaboration, and clear communication are crucial soft skills for success in this role. These skills and qualities are vital for efficiently building, optimizing, and deploying machine learning models in real-world applications.

What is the difference between Pytorch Developer vs Machine Learning Engineer?

AspectPytorch DeveloperMachine Learning Engineer
Required CredentialsBachelor's or higher in CS, experience with PyTorchBachelor's or higher in CS, data science, or related field, with ML experience
Work EnvironmentResearch labs, AI startups, tech companies focusing on deep learningTech companies, finance, healthcare, often involving deployment and scaling ML models
Industry UsagePrimarily in AI research and development teamsAcross industries implementing ML solutions in production

While both roles require knowledge of machine learning and experience with PyTorch, a Pytorch Developer mainly focuses on developing and optimizing deep learning models using PyTorch. A Machine Learning Engineer often has a broader scope, including deploying, maintaining, and scaling ML models across various platforms and industries.

What are popular job titles related to Pytorch Developer jobs in Santa Cruz, CA?

For Pytorch Developer jobs in Santa Cruz, CA, the most frequently searched job titles are:

What job categories do people searching Pytorch Developer jobs in Santa Cruz, CA look for?

The top searched job categories for Pytorch Developer jobs in Santa Cruz, CA are:

What cities near Santa Cruz, CA are hiring for Pytorch Developer jobs?

Cities near Santa Cruz, CA with the most Pytorch Developer job openings:

Principal Software Developer - AI/ML Performance Validation & Systems Testing

Advanced Micro Devices, Inc

San Jose, CA • On-site

$210K/yr

Full-time

Posted 17 days ago


Advanced Micro Devices rating

8.6

Company rating: 8.6 out of 10

Based on 13 frontline employees who took The Breakroom Quiz

26th of 161 rated electronics manufacturers


Job description

ADVANCE YOUR CAREER. ADVANCE THE WORLD.
At AMD, we believe technology has the power to solve the world's most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future.
Whether you're designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger - technology that moves the world forward. Join us and, together, we'll advance your career.
THE ROLE:
We are seeking a Principal Software Engineer to serve as the senior technical leader for ROCm software validation across compute workloads and server-class systems. In this individual-contributor leadership role, you will define how AMD proves ROCm is ready to ship - from unit and component testing, through full-stack workload validation, to multi-node system-level qualification on AMD Instinct™ GPU platforms. You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production.
THE PERSON:
You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production.
KEY RESPONSIBILITIES:
  • Own the end-to-end validation architecture for ROCm - unit, integration, framework, workload, performance, stress, stability, scale-out, and system-level test layers - across multiple GPU generations and server platforms.
  • Define release-qualification gates and exit criteria for ROCm software releases (functional coverage, performance regressions, stability hours, scale targets, RAS criteria) and drive the org to meet them.
  • Architect the test infrastructure - distributed test runners, GitHub Actions / Jenkins / internal CI fleets, hardware lab orchestration, result data lakes, flaky-test detection, bisection automation, and self-service developer pre-submit pipelines.
  • Champion modern, agile quality engineering - shift-left testing, test pyramids, contract testing between layers, hermetic test environments, deterministic reproducers, and continuous validation in trunk.
  • Set the bar for GitHub-based quality workflows - PR gating policy, required checks, code-coverage standards, bug-bash and triage cadences, and disciplined issue management across ROCm/* repositories and partner upstream projects.
  • Lead complex escalation debug - partner with development, hardware, firmware, and customer-facing teams to root-cause the hardest multi-day, multi-node, multi-component failures and convert findings into durable test coverage.
  • Influence the roadmap - work with product management, silicon, platform, and software architecture to ensure validation readiness for next-generation Instinct GPUs and server platforms before tape-in milestones and silicon arrival.
  • Mentor and elevate Senior and Staff validation engineers, SDETs, and SQA leads; raise the technical bar through design review, code review, and written guidance.
  • Represent ROCm validation externally - strategic customer engagements, OEM qualification programs, and open-source community quality initiatives.
  • Lead system-level testing for server nodes - multi-GPU topologies, PCIe/Infinity Fabric/xGMI, BMC/IPMI, thermal/power, firmware interactions, and multi-node fabric (Ethernet/InfiniBand/UALink) bring-up and validation.Drive compute workload validation and characterization - LLM training and inference (PyTorch, vLLM, Triton, JAX), recommender systems, scientific HPC kernels, MLPerf-class benchmarks - establishing reproducible methodology, baselines, and regression tracking.

PREFERRED EXPERIENCE:
  • Software engineering experience in validation, SDET, or quality engineering, including experience leading complex systems validation.
  • Expert Python for test automation and infrastructure; strong C++ for debugging and extending production code.
  • Deep validation expertise in two or more of the following:
    • GPU software stacks (ROCm, CUDA, oneAPI, SYCL)
    • AI/ML frameworks (PyTorch, TensorFlow, JAX, Triton, vLLM)
    • HPC runtimes and communication libraries (MPI, RCCL/NCCL, UCX, Libfabric)
    • Linux kernel, GPU drivers, or accelerator firmware
    • Distributed systems and large-scale cluster software
  • Experience validating multi-GPU, multi-node server platforms, including stress, soak, fault injection, and RAS testing.
  • Experience defining and delivering release qualification programs for hyperscalers, OEMs, or Tier-1 customers.
  • Contributions to validation, CI, or test infrastructure for ROCm, PyTorch, LLVM, Triton, vLLM, or similar open-source projects.
  • Experience leading adoption of agentic AI workflows, including automated testing, AI-driven debugging, MCP, and RAG-based engineering solutions.
  • Experience validating or operating large-scale GPU clusters (256+ GPUs), including fabric bring-up, health monitoring, and diagnostics.
  • Familiarity with AI training, inference, and HPC benchmark methodologies.
  • Experience with performance validation, profiling tools (rocprof, Omniperf, Nsight), and regression analysis.
  • Familiarity with hardware lab automation, including BMC/IPMI/Redfish, PDU control, serial consoles, automated re-imaging, and topology-aware scheduling.
  • Experience supporting validation for pre-silicon, emulation, and first-silicon accelerator bring-up.

ACADEMIC CREDENTIALS:
  • BS/MS/PhD in Computer Science, Computer Engineering, or related discipline (or equivalent demonstrated experience).

LOCATION: San Jose, California
#LI-DR1
#LI-HYBRID
Benefits offered are described: AMD benefits at a glance.
AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants' needs under the respective laws throughout all stages of the recruitment and selection process.
AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD's "Responsible AI Policy" is available here.
This posting is for an existing vacancy.

What Advanced Micro Devices employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom