Ampere
Ampere

3 Ampere Jobs Hiring Near You

Ampere Jobs Information

What are the most popular titles at Ampere?
What are the most popular cities for Ampere jobs?
Infographic showing various job openings at Ampere in the United States as of July 2026, with employment types broken down into 100% Full Time. Highlights an 100% Physical job distribution.
AI Accelerator Software Principal Engineer - Runtime Library

AI Accelerator Software Principal Engineer - Runtime Library

Ampere

Santa Clara, CA • On-site

$158K - $212K/yr

Full-time

Posted 24 days ago


Job description

Job Summary:
Ampere is a semiconductor design company focused on high-performance, energy-efficient AI compute. The AI Accelerator Principal Software Engineer will lead the design and optimization of AI runtime software, enabling deep learning models to run efficiently on Ampere's accelerators.
Responsibilities:
• Build and evolve an AI Runtime Library for Ampere accelerators that supports execution, scheduling, and lifecycle management of deep learning workloads across multiple model types and popular frameworks.
• Own end-to-end acceleration paths, going deep into the full SW/HW stack—including: Inference serving and integration layers, Compiler/runtime interfaces and graph/IR execution flows, Runtime library architecture (APIs, memory management, operators, execution engines), Communication mechanisms and device/host orchestration
• Drive HW/SW co-design and optimization to improve: Throughput (tokens/requests per second), Latency (kernel execution and scheduling efficiency), Memory efficiency (buffering, paging, reuse, caching), Overall compute utilization and scaling behavior
• Contribute to AI co-processor/accelerator software enablement, partnering closely with hardware and systems teams to ensure runtime and kernel strategies match accelerator capabilities and constraints.
• Collaborate cross-functionally to integrate runtime components into Ampere platform stacks, ensuring robust deployment on target environments and consistent performance in production-like workloads.
Qualifications:
Required:
• BS Computer Science, Computer Engineering, Electrical Engineering, or Software Engineering or related technical field & 8 years of related experience; or MS degree & 6 years; or PhD & 3 years
• Proven experience developing user-mode drivers and/or runtime libraries for GPUs or deep learning accelerators in Linux or RTOS environments.
• Strong expertise in C/C++ and systems-level programming (memory, threading, synchronization, performance profiling).
• Demonstrated background in AI framework enablement, with hands-on experience in one or more of: PyTorch (operator/runtime integration, graph execution, correctness/performance work), llama.cpp (inference/runtime execution patterns), ONNX (graph handling, interoperability, execution engines).
• Strong performance engineering skills, including profiling/diagnostics and optimization of execution pipelines, data movement, and compute kernels.
• Ability to operate effectively in a collaborative environment—owning complex components while partnering with compilers, hardware, and platform teams.
Company:
A semiconductor design company leading the future of computing with an innovative approach to CPU design focused on high-performance, energy efficient AI compute. Founded in 2017, the company is headquartered in Santa Clara, USA, with a team of 1001-5000 employees. The company is currently Late Stage.