LLM Inference Engineer
OR · On-site +1
We are specifically seeking an expert in high-performance LLM serving systems and inference optimization. In this role, you will push the boundaries of how large language models are served. What You ...
OR · On-site +1
We are specifically seeking an expert in high-performance LLM serving systems and inference optimization. In this role, you will push the boundaries of how large language models are served. What You ...
OR · On-site +1
We are specifically seeking an expert in high-performance LLM serving systems and inference optimization. In this role, you will push the boundaries of how large language models are served. What You ...
Sunnyvale, CA · On-site
$122K - $168K/yr
We are looking for a Senior Inference Engineer to own inference for real-time multimodal conversational AI. This is a full-stack inference role: you will work across the entire path a model takes ...
Sunnyvale, CA · On-site
$122K - $168K/yr
We are looking for a Senior Inference Engineer to own inference for real-time multimodal conversational AI. This is a full-stack inference role: you will work across the entire path a model takes ...
Santa Clara, CA · On-site
$74 - $97.50/hr
Build inference pipelines with tools like NVIDIA Dynamo, distributing tasks among GPU workers to improve efficiency. * Collaborate with DevOps teams to orchestrate disaggregated inference using ...
Santa Clara, CA · On-site
$74 - $97.50/hr
Build inference pipelines with tools like NVIDIA Dynamo, distributing tasks among GPU workers to improve efficiency. * Collaborate with DevOps teams to orchestrate disaggregated inference using ...
Sunnyvale, CA · On-site
$122K - $168K/yr
We are looking for a Senior Inference Engineer to own inference for real-time multimodal conversational AI. This is a full-stack inference role: you will work across the entire path a model takes ...
Sunnyvale, CA · On-site
$122K - $168K/yr
We are looking for a Senior Inference Engineer to own inference for real-time multimodal conversational AI. This is a full-stack inference role: you will work across the entire path a model takes ...
$145K - $188K/yr
The Inference team is responsible for building and operating highly scalable, reliable production inference services for both serverless and dedicated inference use cases. The work spans platform ...
Quick apply
$145K - $188K/yr
The Inference team is responsible for building and operating highly scalable, reliable production inference services for both serverless and dedicated inference use cases. The work spans platform ...
Eagan, MN · On-site
$106K - $146K/yr
About the Role As a Senior Inference Engineer, AI , responsibilities include/you will: * Within Platform Engineering and Enterprise AI Services, an AI Inference Engineer is responsible for ...
Eagan, MN · On-site
$106K - $146K/yr
About the Role As a Senior Inference Engineer, AI , responsibilities include/you will: * Within Platform Engineering and Enterprise AI Services, an AI Inference Engineer is responsible for ...
About the Role We're hiring an Inference Engineer to advance our mission of building real-time multimodal intelligence. Your Impact * Design and build low latency, scalable, and reliable model ...
About the Role We're hiring an Inference Engineer to advance our mission of building real-time multimodal intelligence. Your Impact * Design and build low latency, scalable, and reliable model ...
Santa Clara, CA · On-site
$74 - $97.50/hr
Build inference pipelines with tools like NVIDIA Dynamo, distributing tasks among GPU workers to improve efficiency. * Collaborate with DevOps teams to orchestrate disaggregated inference using ...
Santa Clara, CA · On-site
$74 - $97.50/hr
Build inference pipelines with tools like NVIDIA Dynamo, distributing tasks among GPU workers to improve efficiency. * Collaborate with DevOps teams to orchestrate disaggregated inference using ...
$151K - $196K/yr
The Inference team is responsible for building and operating highly scalable, reliable production inference services for both serverless and dedicated inference use cases. The work spans platform ...
Quick apply
$151K - $196K/yr
The Inference team is responsible for building and operating highly scalable, reliable production inference services for both serverless and dedicated inference use cases. The work spans platform ...
New York, NY · On-site
$150K - $350K/yr
The Role We are looking for an Inference Runtime Software Engineer to push forward LM Studio's inference stack on-device and in the cloud. You will integrate new inference engines and runtime ...
New York, NY · On-site
$150K - $350K/yr
The Role We are looking for an Inference Runtime Software Engineer to push forward LM Studio's inference stack on-device and in the cloud. You will integrate new inference engines and runtime ...
Manhattan, NY · On-site
$140 - $210/hr
The Role We are looking for an Inference Runtime Software Engineer to push forward LM Studio's inference stack on-device and in the cloud. You will integrate new inference engines and runtime ...
Manhattan, NY · On-site
$140 - $210/hr
The Role We are looking for an Inference Runtime Software Engineer to push forward LM Studio's inference stack on-device and in the cloud. You will integrate new inference engines and runtime ...
Sunnyvale, CA · On-site
$122K - $168K/yr
We are looking for a Senior Inference Engineer to own inference for real-time multimodal conversational AI. This is a full-stack inference role: you will work across the entire path a model takes ...
Sunnyvale, CA · On-site
$122K - $168K/yr
We are looking for a Senior Inference Engineer to own inference for real-time multimodal conversational AI. This is a full-stack inference role: you will work across the entire path a model takes ...
Palo Alto, CA · On-site
$250K - $350K/yr
About the Role We are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products. In this highly technical role, you will operate at the intersection of ...
Palo Alto, CA · On-site
$250K - $350K/yr
About the Role We are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products. In this highly technical role, you will operate at the intersection of ...
San Francisco, CA · On-site
$175 - $250/hr
Global Inference Library Engineer Experience: Senior Level Salary: $175,000 - $250,000 per year Job Details We're looking for an engineer to help build and maintain a high-performance inference ...
San Francisco, CA · On-site
$175 - $250/hr
Global Inference Library Engineer Experience: Senior Level Salary: $175,000 - $250,000 per year Job Details We're looking for an engineer to help build and maintain a high-performance inference ...
Sunnyvale, CA · On-site
$122K - $168K/yr
We are looking for a Senior Inference Engineer to own inference for real-time multimodal conversational AI. This is a full-stack inference role: you will work across the entire path a model takes ...
Sunnyvale, CA · On-site
$122K - $168K/yr
We are looking for a Senior Inference Engineer to own inference for real-time multimodal conversational AI. This is a full-stack inference role: you will work across the entire path a model takes ...
New York, NY · Remote
$150K - $350K/yr
The Role We are looking for an Inference Runtime Software Engineer to push forward LM Studio's inference stack on-device and in the cloud. You will integrate new inference engines and runtime ...
Quick apply
New York, NY · Remote
$150K - $350K/yr
The Role We are looking for an Inference Runtime Software Engineer to push forward LM Studio's inference stack on-device and in the cloud. You will integrate new inference engines and runtime ...
New York, NY · On-site +1
The Role We are looking for an Inference Runtime Software Engineer to push forward LM Studio's inference stack on-device and in the cloud. You will integrate new inference engines and runtime ...
New York, NY · On-site +1
The Role We are looking for an Inference Runtime Software Engineer to push forward LM Studio's inference stack on-device and in the cloud. You will integrate new inference engines and runtime ...
Burlingame, CA · On-site
$110K - $270K/yr
The AI Inference Engineer at Quadric will [1] port AI models to Quadric platform; [2] optimize the model deployment for efficient inference; [3] profile and benchmark the model performance. This ...
Burlingame, CA · On-site
$110K - $270K/yr
The AI Inference Engineer at Quadric will [1] port AI models to Quadric platform; [2] optimize the model deployment for efficient inference; [3] profile and benchmark the model performance. This ...
Palo Alto, CA · On-site
$180K - $440K/yr
We are building the high-performance inference platform that serves Grok to millions of users every day with lightning speed and perfect reliability. * As a Member of Technical Staff - Inference, you ...
Palo Alto, CA · On-site
$180K - $440K/yr
We are building the high-performance inference platform that serves Grok to millions of users every day with lightning speed and perfect reliability. * As a Member of Technical Staff - Inference, you ...
San Francisco, CA · On-site
$180K - $250K/yr
About the Role We're hiring an Inference Engineer to advance our mission of building real-time multimodal intelligence. Your Impact * Design and build low latency, scalable, and reliable model ...
San Francisco, CA · On-site
$180K - $250K/yr
About the Role We're hiring an Inference Engineer to advance our mission of building real-time multimodal intelligence. Your Impact * Design and build low latency, scalable, and reliable model ...

Full-time
Re-posted yesterday
Locations: San Francisco or Remote
About The Role
The NEAR AI team is building decentralized and confidential machine learning infrastructure to enable user-owned AI. Our mission is to build highly scalable and efficient infrastructure for open-source AI at a global scale.
We are specifically seeking an expert in high-performance LLM serving systems and inference optimization. In this role, you will push the boundaries of how large language models are served.
What You'll Be Doing
What We're Looking For
We'd Love If You Have
Please let us know if you require any special requirements for your interview and we'll do our best to accommodate.