Inference

60 Inference Jobs Hiring Near You

We are specifically seeking an expert in high-performance LLM serving systems and inference optimization. In this role, you will push the boundaries of how large language models are served. What You ...

Senior Inference Engineer, AGI

Sunnyvale, CA · On-site

$122K - $168K/yr

We are looking for a Senior Inference Engineer to own inference for real-time multimodal conversational AI. This is a full-stack inference role: you will work across the entire path a model takes ...

Senior Inference Engineer, AGI

Sunnyvale, CA · On-site

$122K - $168K/yr

We are looking for a Senior Inference Engineer to own inference for real-time multimodal conversational AI. This is a full-stack inference role: you will work across the entire path a model takes ...

The Inference team is responsible for building and operating highly scalable, reliable production inference services for both serverless and dedicated inference use cases. The work spans platform ...

Senior Inference Engineer - AI

Eagan, MN · On-site

$106K - $146K/yr

About the Role As a Senior Inference Engineer, AI , responsibilities include/you will: * Within Platform Engineering and Enterprise AI Services, an AI Inference Engineer is responsible for ...

About the Role We're hiring an Inference Engineer to advance our mission of building real-time multimodal intelligence. Your Impact * Design and build low latency, scalable, and reliable model ...

Senior Inference Engineer, AGI

Sunnyvale, CA · On-site

$122K - $168K/yr

We are looking for a Senior Inference Engineer to own inference for real-time multimodal conversational AI. This is a full-stack inference role: you will work across the entire path a model takes ...

About the Role We are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products. In this highly technical role, you will operate at the intersection of ...

Senior Inference Engineer, AGI

Sunnyvale, CA · On-site

$122K - $168K/yr

We are looking for a Senior Inference Engineer to own inference for real-time multimodal conversational AI. This is a full-stack inference role: you will work across the entire path a model takes ...

AI Inference Engineer

Burlingame, CA · On-site

$110K - $270K/yr

The AI Inference Engineer at Quadric will [1] port AI models to Quadric platform; [2] optimize the model deployment for efficient inference; [3] profile and benchmark the model performance. This ...

Inference Engineer

San Francisco, CA · On-site

$180K - $250K/yr

About the Role We're hiring an Inference Engineer to advance our mission of building real-time multimodal intelligence. Your Impact * Design and build low latency, scalable, and reliable model ...

Showing results 41-60

Inference Jobs Information

Infographic showing various job openings at Inference in the United States as of August 2026, with employment types broken down into 100% Full Time. Highlights an 100% Physical job distribution.

LLM Inference Engineer

Near AI

OR • On-site, Remote

Full-time

Re-posted yesterday


Job description

Locations: San Francisco or Remote

About The Role

The NEAR AI team is building decentralized and confidential machine learning infrastructure to enable user-owned AI. Our mission is to build highly scalable and efficient infrastructure for open-source AI at a global scale.

We are specifically seeking an expert in high-performance LLM serving systems and inference optimization. In this role, you will push the boundaries of how large language models are served.

What You'll Be Doing

  • Architect and maintain production high-traffic LLM serving systems.
  • Optimize throughput, latency, and cost for leading open-source LLMs.

What We're Looking For

  • Strong hands-on experience in LLM inference, with expertise debugging and optimizing major inference engines such as SGLang, vLLM, or TensorRT.
  • Deep knowledge of state-of-the-art GPU architectures, and effectively exploit them using PyTorch, Triton, CuTe, CUDA, etc.
  • Proven track record in designing and maintaining end-to-end high-traffic LLM serving systems.
  • Strong problem-solving skills and ability to communicate technical ideas clearly.

We'd Love If You Have

  • Experience with Trusted Execution Environments (TEE).
  • Active contributor to open-source LLM inference engines.

Please let us know if you require any special requirements for your interview and we'll do our best to accommodate.