1

Contract Cuda Developer Jobs (NOW HIRING)

$140 - $210/hr

Design and optimize custom inference stacks: kernel-level work (CUDA/Triton), quantization (GPTQ ... Opportunity to work on impactful client engagements * Long-term contract engagement with potential ...

$120 - $150/hr

Design, implement, and maintain DevOps pipelines for C++ or Python applications, including ... Experience with GPU and CUDA development for performanceโ€‘critical applications. * Experience ...

Eng Sr - SW

Merrimack, NH ยท On-site

$125K - $165K/yr

... contract, this specific position requires US citizenship status. About BAE Systems Electronic ... Experience with CUDA programming * Familiarity with the RF and software-defined radios * Experience ...

Software Developer

Mclean, VA ยท On-site

$86K - $198K/yr

Experience with CUDA and GPU accelerated development * Ability to work with automated testing tools ... as well as contract-specific affordability and organizational requirements. The projected ...

Experience with CUDA and GPU accelerated development * Ability to work with automated testing tools ... as well as contract-specific affordability and organizational requirements. The projected ...

Software Developer

Mclean, VA ยท On-site +1

$86K - $198K/yr

Experience with CUDA and GPU accelerated development * Ability to work with automated testing tools ... as well as contract-specific affordability and organizational requirements. The projected ...

Software Developer

Mclean, VA ยท On-site

$86K - $198K/yr

Experience with CUDA and GPU accelerated development * Ability to work with automated testing tools ... as well as contract-specific affordability and organizational requirements. The projected ...

Showing results 21-40

Contract Cuda Developer information

See salary details

$17

$52

$81

How much do contract cuda developer jobs pay per hour?

As of Sep 7, 2026, the average hourly pay for contract cuda developer in the United States is $52.84, according to ZipRecruiter salary data. Most workers in this role earn between $40.38 and $64.66 per hour, depending on experience, location, and employer.

What is the difference between Contract Cuda Developer vs Contract GPU Programmer?

AspectContract Cuda DeveloperContract GPU Programmer
Required CredentialsProficiency in CUDA, C++, GPU architecture knowledgeProficiency in GPU programming, CUDA, OpenCL, or similar
Work EnvironmentTech companies, research labs, software development firmsGaming, simulation, scientific computing industries
Employer & Industry UsagePrimarily in tech, AI, and high-performance computing sectorsIn industries utilizing GPU acceleration like gaming and scientific research

The Contract Cuda Developer and Contract GPU Programmer roles both require expertise in GPU technologies and CUDA. However, the Contract Cuda Developer typically focuses more on developing and optimizing CUDA-specific applications, while the Contract GPU Programmer may work across various GPU programming frameworks like OpenCL. Both roles are vital in high-performance computing environments, but their specific focus and industry applications differ slightly.

More about Contract Cuda Developer jobs

What cities are hiring for Contract Cuda Developer jobs?

Cities with the most Contract Cuda Developer job openings:

What are the most commonly searched types of Cuda Developer jobs?

The most popular types of Cuda Developer jobs are:

What states have the most Contract Cuda Developer jobs?

States with the most job openings for Contract Cuda Developer jobs include:

What job categories do people searching Contract Cuda Developer jobs look for?

The top searched job categories for Contract Cuda Developer jobs are:

Infographic showing various Contract Cuda Developer job openings in the United States as of August 2026, with employment types broken down into 85% Full Time, 3% Part Time, and 12% Contract. Highlights an 80% Physical, 5% Hybrid, and 15% Remote job distribution, with an average salary of $109,905 per year, or $52.8 per hour.

Hiring: On-prem Platform Engineer at Charlotte, NC (Onsite)

Realtech Services

Charlotte, NC โ€ข On-site

Contractor

Re-posted 2 days ago


Job description


 
 

Position: On-prem Platform Engineer

Location:  Charlotte, NC (Onsite)

Duration: Long Term Contract

No. of positions :: 3

 

Key Skills:

Must-Have Skills (Mandatory Keywords)

LLM Inference & Optimization

  • vLLM, TensorRT-LLM, Triton Inference Server, SGLang
  • Inference optimization techniques:
    • Continuous batching
    • Speculative decoding
    • KV cache / Prefix caching
  • Model optimization:
    • FP8, AWQ, GPTQ

Distributed & GPU Systems

  • Tensor parallelism and large model scaling
  • CUDA, NCCL, GPU architecture
  • GPU partitioning & optimization (MIG)

Kubernetes & ML Serving

  • Kubernetes-based ML serving platforms
  • KServe, OpenShift AI
  • Helm charts, Operators, platform automation

GPU Orchestration

  • Run:AI or similar GPU scheduling/orchestration platforms
  • Multi-tenant GPU workload management

Platform Engineering

  • Experience building internal AI/ML platforms (on-prem or hybrid)
  • Strong automation and system design mindset

Observability & Performance

  • Prometheus, Grafana
  • ML observability (model latency, throughput, drift, resource utilization)
  • Performance benchmarking and tuning

Good to Have / Preferred Skills:

  • Experience with LLMOps / Gen-AI pipelines
  • Exposure to hybrid cloud (on-prem + GCP/Azure integration)
  • Familiarity with Inferentia / alternative accelerators
  • Knowledge of service mesh / networking in GPU clusters
  • Build, configure, and operate on‑prem Kubernetes/OpenShift AI platforms for deploying and serving Gen-AI models and LLM inference workloads.
  • Design and optimize high‑performance inference stacks using vLLM, Tensor RT‑LLM, Triton Inference Server, SGLang, and advanced techniques (continuous batching, speculative decoding, KV caching).
  • Manage GPU orchestration and capacity using Run: AI, MIG, CUDA/NCCL, and tensor parallelism to maximize utilization and throughput.
  • Deploy and operate Kubernetes ML serving frameworks (KServe, Helm, Operators) for scalable, reliable model serving.
  • Drive inference optimization and benchmarking, leveraging FP8, AWQ, GPTQ, and performance tools such as GuideLLM and Locust.
  • Implement observability and ML monitoring using Prometheus, Grafana, Arize AI, ensuring SLA/SLO compliance for Gen-AI services.
  • Collaborate with ML and research teams to onboard new models, tune inference performance, and productionize Gen-AI use cases.