LLM Inference Deployment Engineer
$180K - $240K/yr
About the Role EnCharge AI is seeking an LLM Inference Deployment Engineer to optimize, deploy, and scale large language models (LLMs) for high-performance inference on its energy efficient AI ...
$180K - $240K/yr
About the Role EnCharge AI is seeking an LLM Inference Deployment Engineer to optimize, deploy, and scale large language models (LLMs) for high-performance inference on its energy efficient AI ...
$180K - $240K/yr
About the Role EnCharge AI is seeking an LLM Inference Deployment Engineer to optimize, deploy, and scale large language models (LLMs) for high-performance inference on its energy efficient AI ...
Santa Clara, CA ยท On-site
$74 - $97.50/hr
Build inference pipelines with tools like NVIDIA Dynamo, distributing tasks among GPU workers to improve efficiency. * Collaborate with DevOps teams to orchestrate disaggregated inference using ...
Santa Clara, CA ยท On-site
$74 - $97.50/hr
Build inference pipelines with tools like NVIDIA Dynamo, distributing tasks among GPU workers to improve efficiency. * Collaborate with DevOps teams to orchestrate disaggregated inference using ...
Livingston, NJ ยท On-site
$140K - $182K/yr
The Inference team is responsible for building and operating highly scalable, reliable production inference services for both serverless and dedicated inference use cases. The work spans platform ...
Livingston, NJ ยท On-site
$140K - $182K/yr
The Inference team is responsible for building and operating highly scalable, reliable production inference services for both serverless and dedicated inference use cases. The work spans platform ...
About the Role We're hiring an Inference Engineer to advance our mission of building real-time multimodal intelligence. Your Impact * Design and build low latency, scalable, and reliable model ...
About the Role We're hiring an Inference Engineer to advance our mission of building real-time multimodal intelligence. Your Impact * Design and build low latency, scalable, and reliable model ...
$220K - $320K/yr
About Inference.net Inference.net trains and hosts specialized language models for companies who want frontier-quality AI at a fraction of the cost. The models we train match GPT-5 accuracy but are ...
$220K - $320K/yr
About Inference.net Inference.net trains and hosts specialized language models for companies who want frontier-quality AI at a fraction of the cost. The models we train match GPT-5 accuracy but are ...
NVIDIA is recruiting a Senior Inference Performance Engineer to push NVIDIA's performance limits on large-scale AI inference benchmarks. This position provides an outstanding opportunity to employ ...
NVIDIA is recruiting a Senior Inference Performance Engineer to push NVIDIA's performance limits on large-scale AI inference benchmarks. This position provides an outstanding opportunity to employ ...
New York, NY ยท On-site
$141K - $182K/yr
The Inference team is responsible for building and operating highly scalable, reliable production inference services for both serverless and dedicated inference use cases. The work spans platform ...
Quick apply
New York, NY ยท On-site
$141K - $182K/yr
The Inference team is responsible for building and operating highly scalable, reliable production inference services for both serverless and dedicated inference use cases. The work spans platform ...
$100K - $140K/yr
Filmmaker / Storyteller Inference.net is seeking a Filmmaker / Storyteller to join our team and help define the narrative of building the world's largest distributed GPU cluster. This role combines ...
$100K - $140K/yr
Filmmaker / Storyteller Inference.net is seeking a Filmmaker / Storyteller to join our team and help define the narrative of building the world's largest distributed GPU cluster. This role combines ...
Santa Clara, CA ยท On-site
$74 - $97.50/hr
Build inference pipelines with tools like NVIDIA Dynamo, distributing tasks among GPU workers to improve efficiency. * Collaborate with DevOps teams to orchestrate disaggregated inference using ...
Santa Clara, CA ยท On-site
$74 - $97.50/hr
Build inference pipelines with tools like NVIDIA Dynamo, distributing tasks among GPU workers to improve efficiency. * Collaborate with DevOps teams to orchestrate disaggregated inference using ...
Sunnyvale, CA ยท On-site
$151K - $196K/yr
The Inference team is responsible for building and operating highly scalable, reliable production inference services for both serverless and dedicated inference use cases. The work spans platform ...
Quick apply
Sunnyvale, CA ยท On-site
$151K - $196K/yr
The Inference team is responsible for building and operating highly scalable, reliable production inference services for both serverless and dedicated inference use cases. The work spans platform ...
Seattle, WA ยท On-site
$332K/yr
What you can expect We are looking for an AI Inference Engineer with a solid background in speech recognition and model inference. In this role, you will develop state-of-the-art automatic speech ...
Seattle, WA ยท On-site
$332K/yr
What you can expect We are looking for an AI Inference Engineer with a solid background in speech recognition and model inference. In this role, you will develop state-of-the-art automatic speech ...
Santa Clara, CA ยท On-site
$74 - $97.50/hr
Build inference pipelines with tools like NVIDIA Dynamo, distributing tasks among GPU workers to improve efficiency. * Collaborate with DevOps teams to orchestrate disaggregated inference using ...
Santa Clara, CA ยท On-site
$74 - $97.50/hr
Build inference pipelines with tools like NVIDIA Dynamo, distributing tasks among GPU workers to improve efficiency. * Collaborate with DevOps teams to orchestrate disaggregated inference using ...
Bellevue, WA ยท On-site
$145K - $188K/yr
The Inference team is responsible for building and operating highly scalable, reliable production inference services for both serverless and dedicated inference use cases. The work spans platform ...
Quick apply
Bellevue, WA ยท On-site
$145K - $188K/yr
The Inference team is responsible for building and operating highly scalable, reliable production inference services for both serverless and dedicated inference use cases. The work spans platform ...
NVIDIA is recruiting a Senior Inference Performance Engineer to push NVIDIA's performance limits on large-scale AI inference benchmarks. This position provides an outstanding opportunity to employ ...
NVIDIA is recruiting a Senior Inference Performance Engineer to push NVIDIA's performance limits on large-scale AI inference benchmarks. This position provides an outstanding opportunity to employ ...
Livingston, NJ ยท On-site
$140K - $182K/yr
The Inference team is responsible for building and operating highly scalable, reliable production inference services for both serverless and dedicated inference use cases. The work spans platform ...
Quick apply
Livingston, NJ ยท On-site
$140K - $182K/yr
The Inference team is responsible for building and operating highly scalable, reliable production inference services for both serverless and dedicated inference use cases. The work spans platform ...
Palo Alto, CA ยท On-site
$134K - $161K/yr
About the Role We are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products. In this highly technical role, you will operate at the intersection of ...
Palo Alto, CA ยท On-site
$134K - $161K/yr
About the Role We are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products. In this highly technical role, you will operate at the intersection of ...
Eagan, MN ยท On-site
$106K - $146K/yr
About the Role As a Senior Inference Engineer, AI , responsibilities include/you will: * Within Platform Engineering and Enterprise AI Services, an AI Inference Engineer is responsible for ...
Eagan, MN ยท On-site
$106K - $146K/yr
About the Role As a Senior Inference Engineer, AI , responsibilities include/you will: * Within Platform Engineering and Enterprise AI Services, an AI Inference Engineer is responsible for ...
Manhattan, NY ยท On-site
$142K - $183K/yr
As a Technical Program Manager focused on inference, you will lead complex programs that enhance the delivery and optimization of inference services, ensuring alignment across various teams to meet ...
Manhattan, NY ยท On-site
$142K - $183K/yr
As a Technical Program Manager focused on inference, you will lead complex programs that enhance the delivery and optimization of inference services, ensuring alignment across various teams to meet ...
New York, NY ยท On-site +1
The Role We are looking for an Inference Runtime Software Engineer to push forward LM Studio's inference stack on-device and in the cloud. You will integrate new inference engines and runtime ...
New York, NY ยท On-site +1
The Role We are looking for an Inference Runtime Software Engineer to push forward LM Studio's inference stack on-device and in the cloud. You will integrate new inference engines and runtime ...
New York, NY ยท Remote
$150K - $350K/yr
The Role We are looking for an Inference Runtime Software Engineer to push forward LM Studio's inference stack on-device and in the cloud. You will integrate new inference engines and runtime ...
Quick apply
New York, NY ยท Remote
$150K - $350K/yr
The Role We are looking for an Inference Runtime Software Engineer to push forward LM Studio's inference stack on-device and in the cloud. You will integrate new inference engines and runtime ...

Sourced by ZipRecruiter
Embedded software
11 - 50 Employees
Santa Clara, CA, US
2022