Tensor

60 Tensor Jobs Hiring in California

Founding GPU Compiler Engineer

San Francisco, CA · On-site

$163K - $200K/yr

SF Tensor is a company focused on rethinking the software and infrastructure stack for AI and high-performance computing. They are seeking a Founding GPU Compiler Engineer to build the core ...

SF Tensor is a company focused on revolutionizing AI and high-performance computing by rethinking software and infrastructure. They are seeking a Founding Product Engineer who will take ownership of ...

About the role Own MLIR dialect design and lowering passes for our AI accelerator - defining the high-level tensor IR, async / streaming semantics, and sharded-tensor types that bridge ML frameworks ...

Responsibilities : • Own MLIR dialect design and lowering passes for our AI accelerator -- defining the high-level tensor IR, async / streaming semantics, and sharded-tensor types that bridge ML ...

Deep understanding of tensor compilation, distributed / sharded execution, and async / streaming dataflow models * Strong C++ fluency and experience integrating with ML frameworks (PyTorch, JAX, ONNX ...

Deep understanding of tensor compilation, distributed / sharded execution, and async / streaming dataflow models * Strong C++ fluency and experience integrating with ML frameworks (PyTorch, JAX, ONNX ...

About the role Own verification of our AI compute core - tensor pipelines, MAC arrays, accumulator logic, and the compute memory interconnect. Work with chip-design and software teams driving ...

Lead the definition of mechanisms for efficient movement of tensor activations, weights, and outputs through on-chip and off-chip memory pathways and high-throughput DMA architecture. * Partner ...

next page

Showing results 1-20

Founding GPU Kernel Engineer

SF Tensor

San Francisco, CA • On-site

Full-time

Re-posted 28 days ago


Job description

Job Summary:
SF Tensor is a company focused on revolutionizing AI and high-performance computing through innovative software and infrastructure solutions. They are seeking a Founding GPU Kernel Engineer who will optimize GPU kernels for machine learning workloads and develop automated compiler passes to enhance performance across various GPU architectures.
Responsibilities:
• Write and hand-optimize GPU kernels for ML workloads (matmuls, attention, normalization, etc.) to set the performance ceilings
• Profile at the microarchitectural level: look into SM utilization, warp stalls, memory bank conflicts, register pressure, instruction throughput
• Debug performance issues by digging deep into things like clock speeds, thermal throttling, driver behavior, hardware errata
• Turn your hand-optimization insights into automated compiler passes (working closely with our compiler team)
• Develop performance models that predict how kernels will behave across different GPU architectures
• Build tools and methods for systematic kernel optimization
• Work with NVIDIA, AMD, and emerging AI accelerators - understand the common parts and what's vendor-specific
Qualifications:
Required:
• Deep expertise in GPU architecture
• Proven track record of hand-writing kernels that match or beat vendor libraries (cuBLAS, cuDNN, CUTLASS)
• Strong skills with low-level profiling tools: Nsight Compute, Nsight Systems, rocprof, or equivalents
• Experience reading and reasoning about PTX/SASS or GPU assembly
• Solid systems programming in C++ and CUDA (or ROCm/HIP)
• Good understanding of how high-level ML operations map to hardware execution
• Experience with distributed training systems: collective ops like all-reduce and all-gather, NCCL/RCCL, multi-node communication patterns
Preferred:
• HPC background: experience with large-scale scientific computing, MPI, or work in supercomputing
• Background in electrical engineering, computer architecture, or hardware design
• Driver development experience (NVIDIA, AMD, or other accelerators)
• Experience with MLIR, LLVM, or compiler backends
• Deep knowledge of distributed ML training: gradient accumulation, activation checkpointing, pipeline/tensor parallelism, ZeRO-style optimizations
• Familiarity with custom accelerators: TPUs, Trainium, Inferentia, or similar
• Knowledge of high-speed interconnects: NVLink, NVSwitch, InfiniBand, RoCE
• Publications or contributions in GPU optimization, HPC, or ML systems
• Experience at NVIDIA, AMD, a national lab, or an AI hardware/infrastructure company
Company:
The San Francisco Tensor Company is reinventing the software and infrastructure stack for modern AI and HPC. Founded in 2025, the company is headquartered in San Francisco, USA, with a team of 2-10 employees. The company is currently Early Stage.