1

Contract Cuda Developer Jobs in Austin, TX (NOW HIRING)

Contract Cuda Developer information

See Austin, TX salary details

$16

$52

$81

How much do contract cuda developer jobs pay per hour?

As of Aug 19, 2026, the average hourly pay for contract cuda developer in Austin, TX is $52.37, according to ZipRecruiter salary data. Most workers in this role earn between $40.05 and $64.09 per hour, depending on experience, location, and employer.

What is the difference between Contract Cuda Developer vs Contract GPU Programmer?

AspectContract Cuda DeveloperContract GPU Programmer
Required CredentialsProficiency in CUDA, C++, GPU architecture knowledgeProficiency in GPU programming, CUDA, OpenCL, or similar
Work EnvironmentTech companies, research labs, software development firmsGaming, simulation, scientific computing industries
Employer & Industry UsagePrimarily in tech, AI, and high-performance computing sectorsIn industries utilizing GPU acceleration like gaming and scientific research

The Contract Cuda Developer and Contract GPU Programmer roles both require expertise in GPU technologies and CUDA. However, the Contract Cuda Developer typically focuses more on developing and optimizing CUDA-specific applications, while the Contract GPU Programmer may work across various GPU programming frameworks like OpenCL. Both roles are vital in high-performance computing environments, but their specific focus and industry applications differ slightly.

What are the most commonly searched types of Cuda Developer jobs in Austin, TX?

The most popular types of Cuda Developer jobs in Austin, TX are:

What job categories do people searching Contract Cuda Developer jobs in Austin, TX look for?

The top searched job categories for Contract Cuda Developer jobs in Austin, TX are:

What cities near Austin, TX are hiring for Contract Cuda Developer jobs?

Cities near Austin, TX with the most Contract Cuda Developer job openings:

Infographic showing various Contract Cuda Developer job openings in Austin, TX as of June 2026, with employment types broken down into 3% Internship, 14% As Needed, 34% Full Time, 19% Part Time, 29% Temporary, and 1% Summer. Highlights an 89% Physical, 4% Hybrid, and 7% Remote job distribution, with an average salary of $108,939 per year, or $52.4 per hour.

W2 Role- Machine Learning Performance Engineer - CUDA Python[50% Travel-remote]

SmartIPlace

Austin, TX • Remote

$143K/yr

Contractor

Re-posted 14 days ago


Job description

Title: Machine Learning Performance Engineer - CUDA Python

Work Authorization – USC / GC only

Interview: Video

Duration:  6-month contract maybe extensions  

 

 

Location:  50% travel

 

Duration: 6 month contract

*Must be willing to travel 50% of the time

*Must have strong pre-sales abilities i.e. presentation skills, communication skills, etc.

*Must be willing to help train employees and customers

Your part here is optimizing the performance of our models – both training and inference. We care about efficient large-scale training, low-latency inference in real-time systems, and high-throughput inference in research.

Part of this is improving straightforward CUDA, but the interesting part needs a whole-systems approach, including storage systems, networking, and host- and GPU-level considerations. Zooming in, we also want to ensure our platform makes sense even at the lowest level – is all that throughput actually goodput? Does loading that vector from the L2 cache really take that long?

  • An understanding of modern ML techniques and toolsets
  • The experience and systems knowledge required to debug a training run’s performance end to end
  • Low-level GPU knowledge of PTX, SASS, warps, cooperative groups, Tensor Cores, and the memory hierarchy
  • Debugging and optimization experience using tools like CUDA GDB, NSight Systems, NSight Compute
  • Library knowledge of Triton, CUTLASS, CUB, Thrust, cuDNN, and cuBLAS
  • Intuition about the latency and throughput characteristics of CUDA graph launch, tensor core arithmetic, warp-level synchronization, and asynchronous memory loads
  • Background in Infiniband, RoCE, GPUDirect, PXN, rail optimization, and NVLink, and how to use these networking technologies to link up GPU clusters
  • An understanding of the collective algorithms supporting distributed GPU training in NCCL or MPI
  • An inventive approach and the willingness to ask hard questions about whether we're taking the right approaches and using the right tools

Smart-iPlace logo

About Smart-iPlace

Sourced by ZipRecruiter

SMART-iPLACE provides innovative staffing and consulting solutions that help our clients achieve their business objectives. We can understand and support all areas of your IT systems from back-end infrastructure to front-end personal productivity. Our goal is create innovative IT solutions that enable your business to be more agile and competitive.

Industry

It services

Company size

51 - 200 Employees

Headquarters location

Irving, TX, US

Year founded

2021

Social media