... direct-to-DN, pNFS layouts, S3 via MDN) -- informed by how HPC and hyperscale environments actually push storage systems (checkpointing, small-file metadata storms, GPU-starved read patterns, mixed ...
... direct-to-DN, pNFS layouts, S3 via MDN) -- informed by how HPC and hyperscale environments actually push storage systems (checkpointing, small-file metadata storms, GPU-starved read patterns, mixed ...
Performance Engineer
Santa Clara, CA · On-site
Performance Engineer Preferred Location: Highly desired in Santa Clara , CA , or Costa Rica (Lower ... NFSv3 direct-to-DN, pNFS layouts, S3 via MDN) informed by how HPC and hyperscale environments ...
Performance Engineer
Santa Clara, CA · On-site
Performance Engineer Preferred Location: Highly desired in Santa Clara , CA , or Costa Rica (Lower ... NFSv3 direct-to-DN, pNFS layouts, S3 via MDN) informed by how HPC and hyperscale environments ...
Sr. HPC Cloud Developer Lead
Mountain View, CA · Hybrid
$66.25 - $90.75/hr
RedLine Performance Solutions (RedLine) has been in the HPC solutions engineering services business ... This full-time, direct hire position offers a full benefits package including paid time off, 401k ...
Sr. HPC Cloud Developer Lead
Mountain View, CA · Hybrid
$66.25 - $90.75/hr
RedLine Performance Solutions (RedLine) has been in the HPC solutions engineering services business ... This full-time, direct hire position offers a full benefits package including paid time off, 401k ...
Performance Engineer
Santa Clara, CA · On-site
... NFSv3 direct-to-DN, pNFS layouts, S3 via MDN) informed by how HPC and hyperscale environments ... scale) that engineering and field teams can use to win technical evaluations and POCs. Drive ...
Performance Engineer
Santa Clara, CA · On-site
... NFSv3 direct-to-DN, pNFS layouts, S3 via MDN) informed by how HPC and hyperscale environments ... scale) that engineering and field teams can use to win technical evaluations and POCs. Drive ...
Sr. Software Development Engineer, HPC/ML Networking Engineer, Annapurna Labs
Cupertino, CA · On-site
$128K - $177K/yr
Experience with embedded systems is valued, and experience with high-speed networking or HPC ... The team enjoys working with numerous principal-level engineers and closely with directors, career ...
Sr. Software Development Engineer, HPC/ML Networking Engineer, Annapurna Labs
Cupertino, CA · On-site
$128K - $177K/yr
Experience with embedded systems is valued, and experience with high-speed networking or HPC ... The team enjoys working with numerous principal-level engineers and closely with directors, career ...
Sr. Software Development Engineer, HPC/ML Networking Engineer, Annapurna Labs
Cupertino, CA · On-site
$128K - $177K/yr
Experience with embedded systems is valued, and experience with high-speed networking or HPC ... The team enjoys working with numerous principal-level engineers and closely with directors, career ...
Sr. Software Development Engineer, HPC/ML Networking Engineer, Annapurna Labs
Cupertino, CA · On-site
$128K - $177K/yr
Experience with embedded systems is valued, and experience with high-speed networking or HPC ... The team enjoys working with numerous principal-level engineers and closely with directors, career ...
Sr. Software Development Engineer, HPC/ML Networking Engineer, Annapurna Labs
Cupertino, CA · On-site
$151K - $199K/yr
If you like solving hard problems, want to work with HPC and ML customers, iterate fast and deliver ... The team enjoys working with numerous principal-level engineers and closely with directors, career ...
Sr. Software Development Engineer, HPC/ML Networking Engineer, Annapurna Labs
Cupertino, CA · On-site
$151K - $199K/yr
If you like solving hard problems, want to work with HPC and ML customers, iterate fast and deliver ... The team enjoys working with numerous principal-level engineers and closely with directors, career ...
... direct engagement. • Develop and drive adoption of new technologies and protocols. • Make ... NVIDIA is a computing platform company operating at the intersection of graphics, HPC, and AI.
... direct engagement. • Develop and drive adoption of new technologies and protocols. • Make ... NVIDIA is a computing platform company operating at the intersection of graphics, HPC, and AI.
Principal DevOps Engineer
El Segundo, CA · On-site
$56.25 - $77/hr
... direct exposure to HPC environments, simulation workloads, or academic/research computing (SLURM, PBS, MPI, Lustre, EFS, AWS ParallelCluster, PCS, etc.). • Cloud DevOps experience, production-grade ...
Principal DevOps Engineer
El Segundo, CA · On-site
$56.25 - $77/hr
... direct exposure to HPC environments, simulation workloads, or academic/research computing (SLURM, PBS, MPI, Lustre, EFS, AWS ParallelCluster, PCS, etc.). • Cloud DevOps experience, production-grade ...
... direct engagement. • Develop and drive adoption of new technologies and protocols. • Make ... HPC programming models and libraries (CUDA, cuDNN, DOCA). • Knowledge of enterprise storage ...
... direct engagement. • Develop and drive adoption of new technologies and protocols. • Make ... HPC programming models and libraries (CUDA, cuDNN, DOCA). • Knowledge of enterprise storage ...
Software Engineer, GPU Infrastructure - HPC
San Francisco, CA · On-site
$230K - $490K/yr
About the role As a software engineer on the Fleet High Performance Computing (HPC) team, you will ... have a direct, adverse and negative relationship with the following job duties, potentially ...
Software Engineer, GPU Infrastructure - HPC
San Francisco, CA · On-site
$230K - $490K/yr
About the role As a software engineer on the Fleet High Performance Computing (HPC) team, you will ... have a direct, adverse and negative relationship with the following job duties, potentially ...
Principal Firmware Engineer - Server Manageability and Observability
Santa Clara, CA · On-site
$272 - $431.25/hr
Align NVIDIA's roadmap with major customers' requirements through direct engagement. * Develop and ... Familiarity with NVIDIA HPC programming models and libraries (CUDA, cuDNN, DOCA). * Knowledge of ...
Principal Firmware Engineer - Server Manageability and Observability
Santa Clara, CA · On-site
$272 - $431.25/hr
Align NVIDIA's roadmap with major customers' requirements through direct engagement. * Develop and ... Familiarity with NVIDIA HPC programming models and libraries (CUDA, cuDNN, DOCA). * Knowledge of ...
Research Computing Engineer
Santa Clara, CA · On-site
$115K - $129K/yr
As the strategic anchor for the SCU High-Performance Computing (HPC) environment, the Research ... The ideal candidate is highly curious, creative, tenacious, and entirely self-directed. They bring ...
Research Computing Engineer
Santa Clara, CA · On-site
$115K - $129K/yr
As the strategic anchor for the SCU High-Performance Computing (HPC) environment, the Research ... The ideal candidate is highly curious, creative, tenacious, and entirely self-directed. They bring ...
Research Computing Engineer
Santa Clara, CA · On-site +1
$115K - $129K/yr
As the strategic anchor for the SCU High-Performance Computing (HPC) environment, the Research ... The ideal candidate is highly curious, creative, tenacious, and entirely self-directed. They bring ...
Research Computing Engineer
Santa Clara, CA · On-site +1
$115K - $129K/yr
As the strategic anchor for the SCU High-Performance Computing (HPC) environment, the Research ... The ideal candidate is highly curious, creative, tenacious, and entirely self-directed. They bring ...
Sr. Director, Sales
San Jose, CA · On-site
$165 - $206/hr
... HPC and IoT/Embedded customers worldwide. We are the #5 fastest growing company among the Silicon ... We seek talented, passionate, and committed engineers, technologists, and business leaders to join ...
Sr. Director, Sales
San Jose, CA · On-site
$165 - $206/hr
... HPC and IoT/Embedded customers worldwide. We are the #5 fastest growing company among the Silicon ... We seek talented, passionate, and committed engineers, technologists, and business leaders to join ...
Sr. Director, Sales
San Jose, CA · On-site
... HPC and IoT/Embedded customers worldwide. We are the #5 fastest growing company among the Silicon ... We seek talented, passionate, and committed engineers, technologists, and business leaders to join ...
Sr. Director, Sales
San Jose, CA · On-site
... HPC and IoT/Embedded customers worldwide. We are the #5 fastest growing company among the Silicon ... We seek talented, passionate, and committed engineers, technologists, and business leaders to join ...
Sr. Director, Sales
San Jose, CA · On-site
... HPC and IoT/Embedded customers worldwide. We are the #5 fastest growing company among the Silicon ... We seek talented, passionate, and committed engineers, technologists, and business leaders to join ...
Sr. Director, Sales
San Jose, CA · On-site
... HPC and IoT/Embedded customers worldwide. We are the #5 fastest growing company among the Silicon ... We seek talented, passionate, and committed engineers, technologists, and business leaders to join ...
Senior Fortran Compiler Engineer
Santa Clara, CA · On-site
$122K - $168K/yr
NVIDIA's HPC compiler group is seeking a Fortran compiler developer to contribute to the ... Direct experience with Flang is a huge plus • Experience writing code using Modern C++ • ...
Senior Fortran Compiler Engineer
Santa Clara, CA · On-site
$122K - $168K/yr
NVIDIA's HPC compiler group is seeking a Fortran compiler developer to contribute to the ... Direct experience with Flang is a huge plus • Experience writing code using Modern C++ • ...
Scientific Computing & HPC Platform Engineering * Lead the architecture, build-out, and ongoing optimization of on-premise GPU clusters, hybrid cloud HPC environments, and supporting storage and ...
Scientific Computing & HPC Platform Engineering * Lead the architecture, build-out, and ongoing optimization of on-premise GPU clusters, hybrid cloud HPC environments, and supporting storage and ...
... HPC software stack. We're looking for a strong technical architect to own the end-to-end ... Align NVIDIA's roadmap with major customers' requirements through direct engagement. * Develop and ...
... HPC software stack. We're looking for a strong technical architect to own the end-to-end ... Align NVIDIA's roadmap with major customers' requirements through direct engagement. * Develop and ...
Director Hpc Engineer information

Other
Posted 6 days ago
Job description
Job Title: Performance Engineer
Location: Santa Clara , CA OR Costa Rica
Duration: 12 + Months
Introduction
- FlashBlade//EXA is Pure Storage''''s exascale, disaggregated storage architecture purpose-built for AI factories—separating metadata (MDN) from data (DN) so each scales independently, and putting clients on a direct pNFS data path to the node that owns their bytes. As the Performance Technical Lead, you will own the performance story for //EXA end-to-end: how it''''s engineered, how it''''s proven, and how it''''s positioned against the fastest-moving competitors in HPC and hyperscale AI infrastructure. This is a hybrid technical/strategic role — equal parts systems performance engineer, benchmark architect, and competitive analyst — reporting into the //EXA technical leadership and working directly with engineering, product management, and field/sales engineering.
What you''''ll do
- Set the performance architecture agenda. Bring deep, current expertise across file, block, and object storage protocols and translate it into concrete performance requirements for //EXA''''s data path (NFSv3 direct-to-DN, pNFS layouts, S3 via MDN) — informed by how HPC and hyperscale environments actually push storage systems (checkpointing, small-file metadata storms, GPU-starved read patterns, mixed-tenant burst I/O).
- Track and act on the NeoCloud / sovereign-cloud shift. Maintain a living view of where //EXA''''s highest-value deployments are heading — GPU-cloud and sovereign-cloud operators (CoreWeave, Crusoe, Nscale, and similar) — and make sure //EXA''''s performance roadmap, reference architectures, and sizing guidance map to how these operators actually buy and operate infrastructure (multi-tenant GPU clusters, bursty training/inference mixes, strict SLAs to end customers).
- Own competitive performance positioning. Build and maintain deep, technically substantiated comparisons against VAST Data, DDN, and WEKA — not marketing bullet points, but real architectural analysis (metadata scaling model, erasure coding/durability tradeoffs, protocol support, GPU-direct paths, cost/performance at scale) that engineering and field teams can use to win technical evaluations and POCs.
- Drive performance tuning for multi-tenant HPC/AI workloads. Lead tuning and validation work spanning the full stack a GPU cluster touches — storage (MDN/DN geometry, pack groups, erasure coding layout), networking (RDMA, RoCE/InfiniBand fabric behavior, NIC/queue tuning), and compute (GPU-side I/O patterns, checkpoint/restore, data loader behavior) — with particular focus on how these interact when multiple tenants/workloads share the same //EXA fleet.
- Build and run the benchmark suite. Own //EXA''''s benchmark framework and result credibility: MLPerf Storage (v2/v3), elbencho, IO500, fio/vdbench-class synthetic tests, and workload-representative benchmarks for AI training/inference and traditional HPC. Ensure results are reproducible, defensible in public disclosure, and directly comparable to published competitor numbers.
- Define QoS, limits, and workload segmentation. Drive the technical requirements and validation for quality-of-service guarantees, per-tenant/per-workload throughput and IOPS limits, and workload isolation — the mechanisms that let //EXA make hard SLA commitments in shared, multi-tenant NeoCloud deployments rather than best-effort performance.
What makes you competitive for this role
- Deep, hands-on background in storage performance engineering across file, block, and object protocols, ideally with direct HPC or hyperscale exposure (parallel filesystems, pNFS/NFS at scale, S3-scale object stores).
- Working knowledge of GPU cluster architecture — RDMA fabrics, GPUDirect Storage, checkpoint/restore patterns for large model training — and how storage bottlenecks manifest in mixed compute/network/storage systems.
- Fluency with industry benchmark standards (MLPerf Storage, IO500) and load-generation tooling (elbencho, fio, vdbench), plus the judgment to design workload-representative tests beyond canned benchmarks.
- Demonstrated ability to build rigorous, technically credible competitive analysis (not slideware) against systems like VAST, DDN, and WEKA — architecture-level understanding, not just spec-sheet comparison.
- Experience with multi-tenant resource management concepts (QoS, rate limiting, workload isolation) in a distributed systems context.
- Comfortable operating across the stack and across audiences — deep enough to debug an RDMA queue-pair stall or a metadata hot-partition, articulate enough to brief a NeoCloud customer''''s technical evaluation team.
Why this role matters
- //EXA''''s disaggregated design (independent MDN/DN scaling, direct-to-data-node pNFS) is a real architectural bet against how VAST, DDN, and WEKA scale metadata and data. Whether that bet wins in the market depends on whether performance claims hold up under the exact multi-tenant, GPU-bound, bursty workloads that NeoCloud and sovereign-cloud operators run — and whether we can prove it with numbers that stand up to public scrutiny. This role is the person who makes sure both are true.