Role Overview We're hiring a Software Engineer to own the serving infrastructure that connects Rime's inference engines to the world. This role sits at the intersection of ML systems and cloud ...
Role Overview We're hiring a Software Engineer to own the serving infrastructure that connects Rime's inference engines to the world. This role sits at the intersection of ML systems and cloud ...
They are seeking an ML Model Serving Engineer to enhance their serving layer with various LLM, speech, and vision models, while collaborating with infrastructure and training engineers to optimize ...
They are seeking an ML Model Serving Engineer to enhance their serving layer with various LLM, speech, and vision models, while collaborating with infrastructure and training engineers to optimize ...
Inference Infrastructure Engineer, Serving
Palo Alto, CA · On-site +1
$200K - $400K/yr
Build low-latency, high-throughput inference serving systems for our large multimodal models * Design and implement techniques that improve latency, throughput, and efficiency, including quantization ...
Inference Infrastructure Engineer, Serving
Palo Alto, CA · On-site +1
$200K - $400K/yr
Build low-latency, high-throughput inference serving systems for our large multimodal models * Design and implement techniques that improve latency, throughput, and efficiency, including quantization ...
Databricks' Model Serving product provides enterprises with a unified, scalable, and governed platform to deploy and manage AI/ML models - from traditional ML to fine-tuned and proprietary large ...
Databricks' Model Serving product provides enterprises with a unified, scalable, and governed platform to deploy and manage AI/ML models - from traditional ML to fine-tuned and proprietary large ...
Senior Software Engineer, Model Serving
$144K - $190K/yr
Databricks' Model Serving product provides enterprises with a unified, scalable, and governed platform to deploy and manage AI/ML models - from traditional ML to fine-tuned and proprietary large ...
Senior Software Engineer, Model Serving
$144K - $190K/yr
Databricks' Model Serving product provides enterprises with a unified, scalable, and governed platform to deploy and manage AI/ML models - from traditional ML to fine-tuned and proprietary large ...
The Staff Software Engineer, Model Serving will play a critical role in designing and building systems for high-throughput, low-latency inference, collaborating across teams to deliver a world-class ...
The Staff Software Engineer, Model Serving will play a critical role in designing and building systems for high-throughput, low-latency inference, collaborating across teams to deliver a world-class ...
Staff Software Engineer, Model Serving
San Francisco, CA · On-site
$192K - $260K/yr
Databricks' Model Serving product provides enterprises with a unified, scalable, and governed platform to deploy and manage AI/ML models - from traditional ML to fine-tuned and proprietary large ...
Staff Software Engineer, Model Serving
San Francisco, CA · On-site
$192K - $260K/yr
Databricks' Model Serving product provides enterprises with a unified, scalable, and governed platform to deploy and manage AI/ML models - from traditional ML to fine-tuned and proprietary large ...
Senior Software Engineer, Model Serving
San Francisco, CA · On-site
$144K - $190K/yr
The Senior Software Engineer for Model Serving will design and implement core systems and APIs, ensuring high-performance and reliability for AI/ML model deployment and management. Responsibilities ...
Senior Software Engineer, Model Serving
San Francisco, CA · On-site
$144K - $190K/yr
The Senior Software Engineer for Model Serving will design and implement core systems and APIs, ensuring high-performance and reliability for AI/ML model deployment and management. Responsibilities ...
Senior Software Engineer, Model Serving
San Francisco, CA · On-site
$166K - $225K/yr
Databricks' Model Serving product provides enterprises with a unified, scalable, and governed platform to deploy and manage AI/ML models - from traditional ML to fine-tuned and proprietary large ...
Senior Software Engineer, Model Serving
San Francisco, CA · On-site
$166K - $225K/yr
Databricks' Model Serving product provides enterprises with a unified, scalable, and governed platform to deploy and manage AI/ML models - from traditional ML to fine-tuned and proprietary large ...
Senior AI Serving Engineer, Backend
San Francisco, CA · On-site
$190K - $250K/yr
You'll help build the model serving platform working across C++, Python, runtime execution, and distributed infrastructure to create a fast, reliable engine for real-time AI applications. You'll gain ...
Senior AI Serving Engineer, Backend
San Francisco, CA · On-site
$190K - $250K/yr
You'll help build the model serving platform working across C++, Python, runtime execution, and distributed infrastructure to create a fast, reliable engine for real-time AI applications. You'll gain ...
Responsibilities : • Design and implement core systems and APIs that power Databricks Foundation Model Serving, ensuring scalability, reliability, and operational excellence. • Partner with ...
Responsibilities : • Design and implement core systems and APIs that power Databricks Foundation Model Serving, ensuring scalability, reliability, and operational excellence. • Partner with ...
Lead Software Engineer, Model Serving Platform
San Francisco, CA · On-site
$230K - $300K/yr
Sciforium is an AI infrastructure company developing next-generation multimodal AI models and a proprietary, high-efficiency serving platform. Backed by multi-million-dollar funding and direct ...
Lead Software Engineer, Model Serving Platform
San Francisco, CA · On-site
$230K - $300K/yr
Sciforium is an AI infrastructure company developing next-generation multimodal AI models and a proprietary, high-efficiency serving platform. Backed by multi-million-dollar funding and direct ...
Foundation Model Serving is the API Product for hosting and serving frontier AI model inference for open source models like Llama, Qwen, and GPT OSS as well as proprietary models like Claude and ...
Foundation Model Serving is the API Product for hosting and serving frontier AI model inference for open source models like Llama, Qwen, and GPT OSS as well as proprietary models like Claude and ...
Experience with LLM serving and routing fundamentals (e.g. rate limiting, token streaming, load balancing, budgets, etc.) * Experience with LLM capabilities and concepts such as reasoning, tool ...
Experience with LLM serving and routing fundamentals (e.g. rate limiting, token streaming, load balancing, budgets, etc.) * Experience with LLM capabilities and concepts such as reasoning, tool ...
Staff Software Engineer, Foundational Model Serving
San Francisco, CA · On-site
$192K - $260K/yr
Foundation Model Serving is the API Product for hosting and serving frontier AI model inference for open source models like Llama, Qwen, and GPT OSS as well as proprietary models like Claude and ...
Staff Software Engineer, Foundational Model Serving
San Francisco, CA · On-site
$192K - $260K/yr
Foundation Model Serving is the API Product for hosting and serving frontier AI model inference for open source models like Llama, Qwen, and GPT OSS as well as proprietary models like Claude and ...
Identify ads management and serving needs unique to Live events due to high unpredictability, massive traffic spikes, and simultaneous global delivery. * Take a holistic yield management approach ...
Identify ads management and serving needs unique to Live events due to high unpredictability, massive traffic spikes, and simultaneous global delivery. * Take a holistic yield management approach ...
Caregiver for Serving Our Nations Heroes
Carlsbad, CA · On-site
$22 - $24/hr
We are incredibly proud of our partnership with the VA and our commitment to serving those who have served our country. We are looking for dependable, heart-centered caregivers to join us in ...
Quick apply
Caregiver for Serving Our Nations Heroes
Carlsbad, CA · On-site
$22 - $24/hr
We are incredibly proud of our partnership with the VA and our commitment to serving those who have served our country. We are looking for dependable, heart-centered caregivers to join us in ...
Identify ads management and serving needs unique to Live events due to high unpredictability, massive traffic spikes, and simultaneous global delivery. * Take a holistic yield management approach ...
Identify ads management and serving needs unique to Live events due to high unpredictability, massive traffic spikes, and simultaneous global delivery. * Take a holistic yield management approach ...
The role involves designing and building platforms for scalable serving of large language models (LLMs), collaborating with researchers and engineers, and leading projects to ensure high-performance ...
The role involves designing and building platforms for scalable serving of large language models (LLMs), collaborating with researchers and engineers, and leading projects to ensure high-performance ...
Identify ads management and serving needs unique to Live events due to high unpredictability, massive traffic spikes, and simultaneous global delivery. * Take a holistic yield management approach ...
Identify ads management and serving needs unique to Live events due to high unpredictability, massive traffic spikes, and simultaneous global delivery. * Take a holistic yield management approach ...
Serving information
See California salary details
$5.46 - $7.72
7% of jobs
$7.72 - $9.99
15% of jobs
$10.50 is the 25th percentile. Wages below this are outliers.
$9.99 - $12.25
12% of jobs
The median wage is $13.87 / hr.
$12.25 - $14.51
22% of jobs
$16.40 is the 75th percentile. Wages above this are outliers.
$14.51 - $16.78
22% of jobs
$16.78 - $19.04
11% of jobs
$19.04 - $21.31
4% of jobs
$21.31 - $23.57
2% of jobs
$23.57 - $25.84
2% of jobs
$25.84 - $28.10
1% of jobs
$28.10 - $30.37
1% of jobs
$5
$15
$30
How much do serving jobs pay per hour?
Who gets paid more, host or server?
What is the difference between Serving vs Waitstaff?
| Aspect | Serving | Waitstaff |
|---|---|---|
| Credentials | No formal certifications typically required | No formal certifications typically required |
| Work Environment | Restaurants, cafes, catering events | Restaurants, cafes, catering events |
| Industry Usage | Commonly used in hospitality and food service | Commonly used in hospitality and food service |
| Search & Comparison | Often searched as Serving | Often searched as Waitstaff |
Serving and Waitstaff are terms used interchangeably in the food service industry. Both roles involve taking orders, serving food and beverages, and ensuring customer satisfaction in restaurants and cafes. While 'Serving' is a broader term, 'Waitstaff' specifically refers to the staff members performing these duties. The responsibilities, work environment, and industry usage are very similar, making these terms often interchangeable in job searches and industry discussions.
What jobs pay 4000 a week without a degree?
What are serving jobs?
What are some common challenges servers face during busy shifts and how can they effectively manage them?
Can you make $100,000 as a server?
What are the key skills and qualifications needed to thrive as a Server, and why are they important?
How do you get a serving job?

Full-time
Posted 10 days ago
Job description
Rime is a foundation modeling company that builds voice AI for enterprises running customer experiences at scale. Our models are purpose-built for high-volume conversational deployments, engineered for the accuracy, performance, and deployment flexibility that production environments actually demand.
We started from a different premise than the rest of the field: build voice AI for human connection, not slop. Before we trained a single model, we built our own corpus: full-duplex, studio-quality conversational speech of normal people, recorded and annotated by linguists. It's why our models are unparalleled in naturalism, and it's why enterprises pick Rime when pilots need to make it to production.
Role Overview
We're hiring a Software Engineer to own the serving infrastructure that connects Rime's inference engines to the world. This role sits at the intersection of ML systems and cloud infrastructure — you'll work directly on model inference and cloud infrastructure to build, harden, and scale the systems that stream voice at real-time latency. As Rime moves toward its next-generation architecture, you'll be a core architect of how our models get served.
What You'll Own
-
Architecture and implementation of Rime's TTS serving infrastructure, from GPU-backed inference engines to the API surface.
-
Model optimization from a single-node to disaggregated fleet serving.
-
Compatibility with different NVIDIA hardwares from Hopper to Blackwell and beyond for on-prem and cloud deployments.
-
Continuous integration and deployment workflows for the model serving pipeline.
-
Site reliability: on-call rotation, monitoring, alerting, and observability across the serving stack.
-
Resource provision, cost management across our GPU fleet.
What We're Looking For
-
Hands-on experience with real-time multinode ML serving infrastructure — ML serving framework experience: NVIDIA Dynamo/Triton, vLLM, SGLang, or equivalent.
-
Experience with distributed or disaggregated model serving (Tensor Parallel, Pipeline Parallel, or equivalent).
-
Strong cloud infrastructure fundamentals: Linux internals, networking, containerization (Docker, Kubernetes).
-
IaC experience — Terraform, Packer, or comparable. You should have opinions about how to do this right.
-
On-call is part of the job. You treat production reliability as a shared responsibility.
Nice to Have
-
Experience with multinode training (DDP, FSDP, etc.).
-
Experience with gRPC or other bidirectional binary streaming protocols.
-
Experience with audio streaming and related technologies (WebRTC, WebSockets, etc.).
-
Experience with a multilingual monorepo where you pick the best language out of merit more than personal experience.
-
Experience with multi-cloud infrastructures (AWS, GCP, OCI, etc.).
-
Comfort with configuration management tooling (Ansible, Chef, Puppet, or similar).
-
SRE, DevOps, or platform engineering background at a startup.
-
Experience at an early-stage company.
Why Join Rime
-
Build the serving infrastructure behind a category-defining voice AI company from the ground up.
-
You will bring in experience that no one else currently has at the company: you can help us set the vision.
-
Direct collaboration with the inference, platform, and ML teams — no handoff culture.
-
The systems you build determine what experiences our customers can deploy at scale.
-
Meaningful equity upside at an early stage.
-
High ownership, high standards, low bureaucracy.
-
SF / Bay Area.
At Rime, we...
-
Are outliers
-
Cut through the hype to focus on the craft
-
Move fast with agency and freedom
-
Maintain a growth mindset, finding joy in the struggle
-
Do the right things, knowing that it'll lead to making money
If that sounds like you too, you'll be a great fit for Rime!