SF Tensor is a company focused on revolutionizing AI and high-performance computing through innovative software and infrastructure solutions. They are seeking a Founding GPU Kernel Engineer who will ...
60 Tensor Jobs Hiring in California
SF Tensor is a company focused on revolutionizing AI and high-performance computing through innovative software and infrastructure solutions. They are seeking a Founding GPU Kernel Engineer who will ...
Founding GPU Compiler Engineer
San Francisco, CA · On-site
$163K - $200K/yr
SF Tensor is a company focused on rethinking the software and infrastructure stack for AI and high-performance computing. They are seeking a Founding GPU Compiler Engineer to build the core ...
Founding GPU Compiler Engineer
San Francisco, CA · On-site
$163K - $200K/yr
SF Tensor is a company focused on rethinking the software and infrastructure stack for AI and high-performance computing. They are seeking a Founding GPU Compiler Engineer to build the core ...
Job Summary : SF Tensor is a company focused on AI and high-performance computing, aiming to eliminate bottlenecks in software and infrastructure. The Founding Research Engineer will develop ...
Job Summary : SF Tensor is a company focused on AI and high-performance computing, aiming to eliminate bottlenecks in software and infrastructure. The Founding Research Engineer will develop ...
SF Tensor is a company focused on revolutionizing AI and high-performance computing by rethinking software and infrastructure. They are seeking a Founding Product Engineer who will take ownership of ...
SF Tensor is a company focused on revolutionizing AI and high-performance computing by rethinking software and infrastructure. They are seeking a Founding Product Engineer who will take ownership of ...
About the role Own MLIR dialect design and lowering passes for our AI accelerator - defining the high-level tensor IR, async / streaming semantics, and sharded-tensor types that bridge ML frameworks ...
About the role Own MLIR dialect design and lowering passes for our AI accelerator - defining the high-level tensor IR, async / streaming semantics, and sharded-tensor types that bridge ML frameworks ...
Responsibilities : • Own MLIR dialect design and lowering passes for our AI accelerator -- defining the high-level tensor IR, async / streaming semantics, and sharded-tensor types that bridge ML ...
Responsibilities : • Own MLIR dialect design and lowering passes for our AI accelerator -- defining the high-level tensor IR, async / streaming semantics, and sharded-tensor types that bridge ML ...
Compiler Engineer -- MLIR
Mountain View, CA · On-site
$200K - $360K/yr
Deep understanding of tensor compilation, distributed / sharded execution, and async / streaming dataflow models * Strong C++ fluency and experience integrating with ML frameworks (PyTorch, JAX, ONNX ...
Quick apply
Compiler Engineer -- MLIR
Mountain View, CA · On-site
$200K - $360K/yr
Deep understanding of tensor compilation, distributed / sharded execution, and async / streaming dataflow models * Strong C++ fluency and experience integrating with ML frameworks (PyTorch, JAX, ONNX ...
Compiler Engineer - MLIR
Mountain View, CA · On-site
$200K - $360K/yr
Deep understanding of tensor compilation, distributed / sharded execution, and async / streaming dataflow models * Strong C++ fluency and experience integrating with ML frameworks (PyTorch, JAX, ONNX ...
Compiler Engineer - MLIR
Mountain View, CA · On-site
$200K - $360K/yr
Deep understanding of tensor compilation, distributed / sharded execution, and async / streaming dataflow models * Strong C++ fluency and experience integrating with ML frameworks (PyTorch, JAX, ONNX ...
Responsibilities : • Write and optimize compute kernels for a custom AI accelerator -- tensor operations, data movement patterns, memory hierarchy exploitation • Develop and maintain profiling ...
Responsibilities : • Write and optimize compute kernels for a custom AI accelerator -- tensor operations, data movement patterns, memory hierarchy exploitation • Develop and maintain profiling ...
Write and optimize compute kernels for a custom AI accelerator - tensor operations, data movement patterns, memory hierarchy exploitation * Develop and maintain profiling infrastructure to measure ...
Write and optimize compute kernels for a custom AI accelerator - tensor operations, data movement patterns, memory hierarchy exploitation * Develop and maintain profiling infrastructure to measure ...
AI Core DV Engineer
Mountain View, CA · On-site
About the role Own verification of our AI compute core - tensor pipelines, MAC arrays, accumulator logic, and the compute memory interconnect. Work with chip-design and software teams driving ...
AI Core DV Engineer
Mountain View, CA · On-site
About the role Own verification of our AI compute core - tensor pipelines, MAC arrays, accumulator logic, and the compute memory interconnect. Work with chip-design and software teams driving ...
Senior Software Engineer - Edge AI/GenAI & Multimedia
San Diego, CA · On-site
$130K - $171K/yr
Strong ability to handle tensor pre-processing and post-processing and integrate AI models into end-to-end pipelines across inputs such as camera, audio, and text. • Strong background in Linux ...
Senior Software Engineer - Edge AI/GenAI & Multimedia
San Diego, CA · On-site
$130K - $171K/yr
Strong ability to handle tensor pre-processing and post-processing and integrate AI models into end-to-end pipelines across inputs such as camera, audio, and text. • Strong background in Linux ...
Kernel Engineer (Compute / Accelerator)
Mountain View, CA · On-site
$260K - $320K/yr
Write and optimize compute kernels for a custom AI accelerator -- tensor operations, data movement patterns, memory hierarchy exploitation * Develop and maintain profiling infrastructure to measure ...
Quick apply
Kernel Engineer (Compute / Accelerator)
Mountain View, CA · On-site
$260K - $320K/yr
Write and optimize compute kernels for a custom AI accelerator -- tensor operations, data movement patterns, memory hierarchy exploitation * Develop and maintain profiling infrastructure to measure ...
Staff Embedded Software Engineer - Edge AI/GenAI & Multimedia
San Diego, CA · On-site
$139K - $183K/yr
Strong ability to handle tensor pre-processing and post-processing and integrate AI models into end-to-end pipelines across inputs such as camera, audio, and text. * Strong background in Linux system ...
Staff Embedded Software Engineer - Edge AI/GenAI & Multimedia
San Diego, CA · On-site
$139K - $183K/yr
Strong ability to handle tensor pre-processing and post-processing and integrate AI models into end-to-end pipelines across inputs such as camera, audio, and text. * Strong background in Linux system ...
Kernel Engineer (Compute / Accelerator)
Mountain View, CA · On-site
$260K - $320K/yr
Write and optimize compute kernels for a custom AI accelerator - tensor operations, data movement patterns, memory hierarchy exploitation * Develop and maintain profiling infrastructure to measure ...
Kernel Engineer (Compute / Accelerator)
Mountain View, CA · On-site
$260K - $320K/yr
Write and optimize compute kernels for a custom AI accelerator - tensor operations, data movement patterns, memory hierarchy exploitation * Develop and maintain profiling infrastructure to measure ...
Senior Software Engineer - Edge AI/GenAI & Multimedia
San Diego, CA · On-site
$130K - $171K/yr
Strong ability to handle tensor pre-processing and post-processing and integrate AI models into end-to-end pipelines across inputs such as camera, audio, and text. * Strong background in Linux system ...
Senior Software Engineer - Edge AI/GenAI & Multimedia
San Diego, CA · On-site
$130K - $171K/yr
Strong ability to handle tensor pre-processing and post-processing and integrate AI models into end-to-end pipelines across inputs such as camera, audio, and text. * Strong background in Linux system ...
Lead the definition of mechanisms for efficient movement of tensor activations, weights, and outputs through on-chip and off-chip memory pathways and high-throughput DMA architecture. * Partner ...
Lead the definition of mechanisms for efficient movement of tensor activations, weights, and outputs through on-chip and off-chip memory pathways and high-throughput DMA architecture. * Partner ...
Senior Software Engineer, CUTLASS Platform
Santa Clara, CA · On-site
$143K - $189K/yr
If you are passionate about designing abstractions for Tensor Core and related GPU hardware features in MLIR, Python, and C++ that enable writing high performance kernels, apply to join the CUTLASS ...
Senior Software Engineer, CUTLASS Platform
Santa Clara, CA · On-site
$143K - $189K/yr
If you are passionate about designing abstractions for Tensor Core and related GPU hardware features in MLIR, Python, and C++ that enable writing high performance kernels, apply to join the CUTLASS ...
Senior Software Engineer -- cuEquivariance
Santa Clara, CA · On-site
$143K - $189K/yr
Responsibilities : • Build, implement, and optimize CUDA kernels for equivariant neural network primitives -- tensor products, segmented polynomials, and triangle-based operations -- targeting peak ...
Senior Software Engineer -- cuEquivariance
Santa Clara, CA · On-site
$143K - $189K/yr
Responsibilities : • Build, implement, and optimize CUDA kernels for equivariant neural network primitives -- tensor products, segmented polynomials, and triangle-based operations -- targeting peak ...
Staff Embedded Software Engineer - Edge AI/GenAI & Multimedia
San Diego, CA · On-site
$139K - $183K/yr
Strong ability to handle tensor pre-processing and post-processing and integrate AI models into end-to-end pipelines across inputs such as camera, audio, and text. * Strong background in Linux system ...
Staff Embedded Software Engineer - Edge AI/GenAI & Multimedia
San Diego, CA · On-site
$139K - $183K/yr
Strong ability to handle tensor pre-processing and post-processing and integrate AI models into end-to-end pipelines across inputs such as camera, audio, and text. * Strong background in Linux system ...
Full-time
Re-posted 28 days ago
Job description
SF Tensor is a company focused on revolutionizing AI and high-performance computing through innovative software and infrastructure solutions. They are seeking a Founding GPU Kernel Engineer who will optimize GPU kernels for machine learning workloads and develop automated compiler passes to enhance performance across various GPU architectures.
Responsibilities:
• Write and hand-optimize GPU kernels for ML workloads (matmuls, attention, normalization, etc.) to set the performance ceilings
• Profile at the microarchitectural level: look into SM utilization, warp stalls, memory bank conflicts, register pressure, instruction throughput
• Debug performance issues by digging deep into things like clock speeds, thermal throttling, driver behavior, hardware errata
• Turn your hand-optimization insights into automated compiler passes (working closely with our compiler team)
• Develop performance models that predict how kernels will behave across different GPU architectures
• Build tools and methods for systematic kernel optimization
• Work with NVIDIA, AMD, and emerging AI accelerators - understand the common parts and what's vendor-specific
Qualifications:
Required:
• Deep expertise in GPU architecture
• Proven track record of hand-writing kernels that match or beat vendor libraries (cuBLAS, cuDNN, CUTLASS)
• Strong skills with low-level profiling tools: Nsight Compute, Nsight Systems, rocprof, or equivalents
• Experience reading and reasoning about PTX/SASS or GPU assembly
• Solid systems programming in C++ and CUDA (or ROCm/HIP)
• Good understanding of how high-level ML operations map to hardware execution
• Experience with distributed training systems: collective ops like all-reduce and all-gather, NCCL/RCCL, multi-node communication patterns
Preferred:
• HPC background: experience with large-scale scientific computing, MPI, or work in supercomputing
• Background in electrical engineering, computer architecture, or hardware design
• Driver development experience (NVIDIA, AMD, or other accelerators)
• Experience with MLIR, LLVM, or compiler backends
• Deep knowledge of distributed ML training: gradient accumulation, activation checkpointing, pipeline/tensor parallelism, ZeRO-style optimizations
• Familiarity with custom accelerators: TPUs, Trainium, Inferentia, or similar
• Knowledge of high-speed interconnects: NVLink, NVSwitch, InfiniBand, RoCE
• Publications or contributions in GPU optimization, HPC, or ML systems
• Experience at NVIDIA, AMD, a national lab, or an AI hardware/infrastructure company
Company:
The San Francisco Tensor Company is reinventing the software and infrastructure stack for modern AI and HPC. Founded in 2025, the company is headquartered in San Francisco, USA, with a team of 2-10 employees. The company is currently Early Stage.