Luminal

2 jobs near Columbus, OH

Cloud Inference Engineer

San Francisco, CA · On-site

$65.75 - $87.75/hr

Luminal is a company that builds an AI compiler and serving stack to enhance model performance. They are seeking a Cloud Inference Engineer to deploy and tune models for low latency and high ...

Senior Compiler Engineer

San Francisco, CA · On-site

$123K - $169K/yr

Luminal is focused on optimizing AI models to accelerate and simplify model deployment. They are seeking a Senior Compiler Engineer to design and build core compiler infrastructure, develop backend ...

Cloud Inference Engineer

Luminal

San Francisco, CA • On-site

$65.75 - $87.75/hr

Full-time

Re-posted 24 days ago


Job description

Job Summary:
Luminal is a company that builds an AI compiler and serving stack to enhance model performance. They are seeking a Cloud Inference Engineer to deploy and tune models for low latency and high throughput on Luminal Cloud.
Responsibilities:
• Deploy and tune models with optimizations like KV caching, paged attention, sequence packing, etc.
• Conducting model performance reviews
• Improve scheduler, batcher, autoscaling; profile latency, cost, utilization
• Sometimes write kernels and, yes, occasional tasteful shitposting
Qualifications:
Required:
• CUDA + GPU inference optimization
• vLLM, SGLang, or TensorRT-LLM experience
• KV caching, paged attention, batching, token streaming, etc.
• Distributed compute (with GPUs is a super plus)
• No degree required
Company:
Luminal provides a machine learning compiler and serverless cloud platform that automates PyTorch model optimization and deployment. Founded in 2025, the company is headquartered in San Francisco, USA, with a team of 11-50 employees. The company is currently Early Stage.