Inference

60 Inference Jobs Hiring Near You

Technical Program Manager, Inference

Manhattan, NY · On-site

$142K - $183K/yr

As a Technical Program Manager focused on inference, you will lead complex programs that enhance the delivery and optimization of inference services, ensuring alignment across various teams to meet ...

Senior Inference Engineer - AI

Eagan, MN · Hybrid

$106K - $146K/yr

About the Role As a Senior Inference Engineer, AI , responsibilities include/you will: * Within Platform Engineering and Enterprise AI Services, an AI Inference Engineer is responsible for ...

By developing and evolving high-performance inference infrastructure, you will enable researchers to explore new ideas with a clear understanding of their computational and systems implications. This ...

Filmmaker / Storyteller

San Francisco, CA · On-site

$100K - $140K/yr

Filmmaker / Storyteller Inference.net is seeking a Filmmaker / Storyteller to join our team and help define the narrative of building the world's largest distributed GPU cluster. This role combines ...

Technical Program Manager, Inference

Livingston, NJ · On-site

$140K - $182K/yr

The Technical Program Manager for Inference will lead cross-functional programs to enhance the delivery and operational excellence of AI inference services, ensuring reliability and scalability for ...

Implement inference-time acceleration techniques such as speculative decoding, tree search, KV cache sharing, etc. * Implement distributed networking primitives for efficient multi-server inference ...

The Inference Software Engineer will support porting state-of-the-art models, enhance runtime capabilities, and optimize communication layers within the system. Responsibilities : • Support porting ...

Inference Engineer

San Francisco, CA · On-site

$180K - $250K/yr

About the Role We're hiring an Inference Engineer to advance our mission of building real-time multimodal intelligence. Your Impact * Design and build low latency, scalable, and reliable model ...

The AI Inference Engineer at Quadric will [1] port AI models to Quadric platform; [2] optimize the model deployment for efficient inference; [3] profile and benchmark the model performance. This ...

Implement inference-time acceleration techniques such as speculative decoding, tree search, KV cache sharing, etc. * Implement distributed networking primitives for efficient multi-server inference ...

Showing results 41-60

Inference Jobs Information

Infographic showing various job openings at Inference in the United States as of August 2026, with employment types broken down into 100% Full Time. Highlights an 100% Physical job distribution.

Technical Program Manager, Inference

CoreWeave

Manhattan, NY • On-site

$142K - $183K/yr

Full-time

Re-posted 7 days ago


CoreWeave rating

9.8

Company rating: 9.8 out of 10

Based on 7 frontline employees who took The Breakroom Quiz

1st of 223 rated it services


Job description

Job Summary:
CoreWeave is The Essential Cloud for AI™, delivering a platform that enables innovators to build and scale AI with confidence. As a Technical Program Manager focused on inference, you will lead complex programs that enhance the delivery and optimization of inference services, ensuring alignment across various teams to meet customer needs.
Responsibilities:
• Drive end-to-end program management for inference platform initiatives spanning reliability, customer onboarding, launch readiness, and runtime optimization
• Lead cross-functional programs for customer onboarding across dedicated and serverless inference offerings, ensuring clear ownership, launch criteria, and readiness for strategic customer use cases
• Drive launch readiness for new inference capabilities by aligning teams around real customer outcomes, supportability, and end-to-end validation
• Partner with engineering and product to define and deliver roadmap outcomes for latency, throughput, uptime, operational quality, and price-performance
• Coordinate multi-team execution across platform, infrastructure, and customer-facing teams to deliver reliable and scalable inference services
• Build and operationalize success metrics, dashboards, launch gates, and review cadences to measure service reliability, onboarding readiness, efficiency, and quality across the inference stack
• Establish repeatable processes for release validation, performance regression tracking, launch management, and postmortem follow-through
• Help unify operational processes, support mechanisms, and execution visibility across inference deployment models and customer onboarding paths
• Create strong communication channels between Engineering, Product, Infrastructure, and Go-to-Market teams to align priorities and deliver predictable, high-impact outcomes
Qualifications:
Required:
• Bachelor's degree in a technical field or equivalent practical experience
• 8+ years of technical program management experience in distributed systems, cloud infrastructure, or AI/ML platform engineering
• Proven experience driving large-scale infrastructure or platform programs from concept to production in complex, cross-functional environments
• Strong technical fluency in distributed inference systems, GPU compute, cloud-native architectures, and performance optimization
• Demonstrated success driving measurable improvements in reliability, performance, operational readiness, or customer delivery
• Excellent written and verbal communication skills, with the ability to align engineering, product, infrastructure, and customer-facing stakeholders around shared goals
• Experience with inference-serving systems, model onboarding workflows, rollout strategies, and observability tooling
• Familiarity with launch readiness, supportability, incident follow-through, and release validation for production infrastructure or platform services
• Understanding of customer onboarding for technical products, especially where platform capabilities, infrastructure readiness, and support processes must align for launch
• Experience operating in high-growth environments where roadmap execution, reliability expectations, and customer commitments must be managed in parallel
Company:
CoreWeave provides cloud infrastructure services designed to support artificial intelligence and high-performance computing workloads. Founded in 2017, the company is headquartered in Livingston, USA, with a team of 1001-5000 employees. The company is currently Late Stage.

What CoreWeave employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom