1

Cpp Software Engineer Jobs (NOW HIRING)

$163K - $255K/yr

People. Joining CPP Investments means joining one of the world's most admired and respected ... As a Lead Software Engineer, you will guide and participate in building and supporting value-add ...

Sr. Staff Software Engineer

San Diego, CA · On-site

$130K - $171K/yr

Engineering Group, Engineering Group > Software Engineering General Summary: As a leading ... CPP, with solid understanding of AI fundamentals, model architectures, tensor layouts, tensor ...

next page

Showing results 1-20

Cpp Software Engineer information

See salary details

$63.5K

$147.5K

$205.5K

How much do cpp software engineer jobs pay per year?

As of Sep 10, 2026, the average yearly pay for cpp software engineer in the United States is $147,524.00, according to ZipRecruiter salary data. Most workers in this role earn between $120,000.00 and $173,000.00 per year, depending on experience, location, and employer.

What cities are hiring for Cpp Software Engineer jobs?

Cities with the most Cpp Software Engineer job openings:

What are popular job titles related to Cpp Software Engineer jobs?

For Cpp Software Engineer jobs, the most frequently searched job titles are:

Software Engineer, Inference Runtime

New York, NY • Remote

$150K - $350K/yr

Full-time

Medical, Dental, Vision, PTO

Re-posted 3 days ago


Job description

LM Studio is used by millions of people around the world to run AI on their own computers, and now with Bionic - also in the cloud. Our values prioritize putting the human in the center, and creating tools that we want to use ourselves, and recommend to our friends and family.

As a team, we work with high technical intensity and personal responsibility. We are looking for curious, self-motivated, creative, and technically excellent teammates to join us and build the future of human-AI interactions in software.

The Role

We are looking for an Inference Runtime Software Engineer to push forward LM Studio's inference stack on-device and in the cloud. You will integrate new inference engines and runtime capabilities, bring up new open-weight models and modalities, and optimize model execution for a wide range of CPU and GPU targets. You will also contribute improvements to the open-source projects we build on.

Qualifications

  • Significant experience building production ML systems, inference runtimes, or performance-sensitive infrastructure

  • Strong programming ability in Python and C++

  • Deep understanding of transformer architectures and the mechanics of model inference

  • Experience profiling CPU or GPU workloads and reasoning about compute, memory, synchronization, and data movement

  • Experience with PyTorch and inference systems such as llama.cpp, MLX, ExecuTorch, vLLM, SGLang, or TensorRT-LLM

  • Strong debugging instincts across model code, runtime internals, operating systems, and CPU or GPU execution

  • Takes personal responsibility for the correctness and performance of their work

Bonus Qualifications

  • Past contributions to open-source inference runtime projects such as llama.cpp, MLX, ExecuTorch, vLLM, SGLang, or TensorRT-LLM

Responsibilities

  • Maintain and push forward our inference stack on-device and in the cloud

  • Bring up new model architectures and multimodal models

  • Improve latency, throughput, memory use, and reliability across CPU, CUDA, Metal, Vulkan, and ROCm runtimes

  • Build runtime capabilities for model loading, batching, scheduling, caching, and distributed execution

  • Benchmark and diagnose correctness and performance problems across the inference stack

  • Contribute upstream to open-source projects such as llama.cpp and MLX

Benefits

  • Competitive salary and equity grants

  • Great medical, vision, dental healthcare plans

  • Catered team lunch / expensed dinners in the office

  • Flexible PTO

  • Flexible WFH

  • Sun-drenched office in SoHo in NYC

Compensation Range: $150K - $350K