We are seeking a technical Product Manager to lead BioNeMo Inference. You will define how ... Define requirements for inference optimization: batching, throughput, latency, GPU utilization ...
We are seeking a technical Product Manager to lead BioNeMo Inference. You will define how ... Define requirements for inference optimization: batching, throughput, latency, GPU utilization ...
We are seeking a technical Product Manager to lead BioNeMo Inference. You will define how ... Define requirements for inference optimization: batching, throughput, latency, GPU utilization ...
We are seeking a technical Product Manager to lead BioNeMo Inference. You will define how ... Define requirements for inference optimization: batching, throughput, latency, GPU utilization ...
We are seeking a technical Product Manager to lead BioNeMo Inference. You will define how ... Define requirements for inference optimization: batching, throughput, latency, GPU utilization ...
We are seeking a technical Product Manager to lead BioNeMo Inference. You will define how ... Define requirements for inference optimization: batching, throughput, latency, GPU utilization ...
Responsibilities Improve throughput, cost, and tail latency through KV-cache management, continuous batching, speculative decoding, and quantization. Optimize long-context prefill and decode ...
Responsibilities Improve throughput, cost, and tail latency through KV-cache management, continuous batching, speculative decoding, and quantization. Optimize long-context prefill and decode ...
Improve throughput, cost, and tail latency through KV-cache management, continuous batching, speculative decoding, and quantization. * Optimize long-context prefill and decode workloads based on real ...
Improve throughput, cost, and tail latency through KV-cache management, continuous batching, speculative decoding, and quantization. * Optimize long-context prefill and decode workloads based on real ...
Food Prep
$20 - $23/hr
Juicing and batching fresh juices, straining and bottling to spec * Building and executing daily/weekly prep lists based on projected volume, in coordination with management Recipe & Quality Control
Food Prep
$20 - $23/hr
Juicing and batching fresh juices, straining and bottling to spec * Building and executing daily/weekly prep lists based on projected volume, in coordination with management Recipe & Quality Control
Claims Clerk Supervisor
Long Beach, CA · On-site
$80K/yr
Oversee claims batching, filing of data-entered claims, and daily retrieval of health-plan files (e.g., pulling claims from the Humana SFTP/FTP and health-plan portals). * Manage document imaging in ...
Quick apply
Claims Clerk Supervisor
Long Beach, CA · On-site
$80K/yr
Oversee claims batching, filing of data-entered claims, and daily retrieval of health-plan files (e.g., pulling claims from the Humana SFTP/FTP and health-plan portals). * Manage document imaging in ...
Food Prep
Mill Valley, CA · On-site
$20 - $23/hr
Juicing and batching fresh juices, straining and bottling to spec * Building and executing daily/weekly prep lists based on projected volume, in coordination with management Recipe & Quality Control
Quick apply
Food Prep
Mill Valley, CA · On-site
$20 - $23/hr
Juicing and batching fresh juices, straining and bottling to spec * Building and executing daily/weekly prep lists based on projected volume, in coordination with management Recipe & Quality Control
Plant Manager
Vernon, CA · On-site
$28 - $31/hr
Operate and maintain batching equipment to produce ready mix concrete per specifications. * Coordinate with drivers and dispatch to ensure timely delivery. * Perform daily inspections and maintenance ...
Plant Manager
Vernon, CA · On-site
$28 - $31/hr
Operate and maintain batching equipment to produce ready mix concrete per specifications. * Coordinate with drivers and dispatch to ensure timely delivery. * Perform daily inspections and maintenance ...
Member of Technical Staff, Backend
San Francisco, CA · On-site
$200K - $300K/yr
... management, roles/permissions, and admin/internal tooling * Familiarity with LLM product infrastructure patterns (caching, batching, evals, integrations) is a plus * Data driven orientation: Be able ...
Member of Technical Staff, Backend
San Francisco, CA · On-site
$200K - $300K/yr
... management, roles/permissions, and admin/internal tooling * Familiarity with LLM product infrastructure patterns (caching, batching, evals, integrations) is a plus * Data driven orientation: Be able ...
Director of Product Management
San Francisco, CA · On-site
$274K - $287K/yr
As Director of Product Management, you will define how FriendliAI's inference platform serves ... Partner closely with engineering and research to shape capabilities across routing, batching ...
Director of Product Management
San Francisco, CA · On-site
$274K - $287K/yr
As Director of Product Management, you will define how FriendliAI's inference platform serves ... Partner closely with engineering and research to shape capabilities across routing, batching ...
Member of Technical Staff, Inference & RL Systems
San Francisco, CA · On-site
$225K - $550K/yr
Optimize KV-cache management, batching strategies, and scheduling * Improve throughput and latency for long-context workloads * Build and maintain distributed RL and post-training infrastructure
Member of Technical Staff, Inference & RL Systems
San Francisco, CA · On-site
$225K - $550K/yr
Optimize KV-cache management, batching strategies, and scheduling * Improve throughput and latency for long-context workloads * Build and maintain distributed RL and post-training infrastructure
Experience composing multi-model pipelines and serving them behind APIs - orchestration, batching, autoscaling, and version management. * A trackrecordbuilding the observability,monitoring, and ...
Experience composing multi-model pipelines and serving them behind APIs - orchestration, batching, autoscaling, and version management. * A trackrecordbuilding the observability,monitoring, and ...
Senior Machine Learning Engineer, Services/MLOps
San Jose, CA · On-site
$122K - $168K/yr
Experience composing multi-model pipelines and serving them behind APIs - orchestration, batching, autoscaling, and version management. * A track record building the observability, monitoring, and ...
Senior Machine Learning Engineer, Services/MLOps
San Jose, CA · On-site
$122K - $168K/yr
Experience composing multi-model pipelines and serving them behind APIs - orchestration, batching, autoscaling, and version management. * A track record building the observability, monitoring, and ...
Claims Clerk Supervisor
Long Beach, CA · On-site
$80K/yr
Oversee claims batching, filing of data-entered claims, and daily retrieval of health-plan files (e.g., pulling claims from the Humana SFTP/FTP and health-plan portals). * Manage document imaging in ...
New
Claims Clerk Supervisor
Long Beach, CA · On-site
$80K/yr
Oversee claims batching, filing of data-entered claims, and daily retrieval of health-plan files (e.g., pulling claims from the Humana SFTP/FTP and health-plan portals). * Manage document imaging in ...
New
Senior Machine Learning Engineer, Services/MLOps
$122K - $168K/yr
Experience composing multi-model pipelines and serving them behind APIs - orchestration, batching, autoscaling, and version management. * A trackrecordbuilding the observability,monitoring, and ...
Senior Machine Learning Engineer, Services/MLOps
$122K - $168K/yr
Experience composing multi-model pipelines and serving them behind APIs - orchestration, batching, autoscaling, and version management. * A trackrecordbuilding the observability,monitoring, and ...
Architect scalable routing, batching, scheduling, memory management, and dynamic loading mechanisms for inference workloads * Ensure reliability, reproducibility, and fault tolerance in the inference ...
Architect scalable routing, batching, scheduling, memory management, and dynamic loading mechanisms for inference workloads * Ensure reliability, reproducibility, and fault tolerance in the inference ...
Our Bar Manager should be equally comfortable behind the bar during a busy service, training a ... Oversee batching, prep, garnishes, and quality control * Maintain a clean, organized, and ...
Quick apply
Our Bar Manager should be equally comfortable behind the bar during a busy service, training a ... Oversee batching, prep, garnishes, and quality control * Maintain a clean, organized, and ...
Our Bar Manager should be equally comfortable behind the bar during a busy service, training a ... Oversee batching, prep, garnishes, and quality control * Maintain a clean, organized, and ...
Our Bar Manager should be equally comfortable behind the bar during a busy service, training a ... Oversee batching, prep, garnishes, and quality control * Maintain a clean, organized, and ...
Deep working knowledge of how AI workloads execute at runtime: serving frameworks, batching strategies, GPU memory management, and the performance levers that determine throughput and latency at ...
Quick apply
Deep working knowledge of how AI workloads execute at runtime: serving frameworks, batching strategies, GPU memory management, and the performance levers that determine throughput and latency at ...
Batching Manager information

Product Manager - BioNeMo Inference
Santa Clara, CA • On-site
Other
Re-posted 10 days ago
Nvidia rating
9.6
Based on 18 frontline employees who took The Breakroom Quiz
Job description
NVIDIA is advancing the frontier of AI for biology with BioNeMo, bringing accelerated computing and generative AI to biomolecular research and drug discovery. We are seeking a technical Product Manager to lead BioNeMo Inference. You will define how developers, researchers, and enterprise platform teams deploy, operate, and scale biomolecular AI inference workloads. This role sits at the intersection of AI infrastructure, developer experience, and product execution: translating the needs of model developers and end users into simple, reliable inference products built on NVIDIA NIM and accelerated computing. A biology or healthcare background is not required. We are looking for a strong technical PM who understands the fundamentals of AI inference serving and model deployment and is eager to apply them to a new and high-impact domain.
What you’ll be doing:- Define product vision, strategy, and roadmap for BioNeMo inference products, including NIM-based deployment, performance optimization, scalability, and developer onboarding.
- Work closely with engineering, research, solution architects, cloud, and product teams to translate model capabilities into production-ready inference experiences.
- Define requirements for inference optimization: batching, throughput, latency, GPU utilization, multi-GPU and multi-node scaling, caching, scheduling, observability, and reliability.
- Partner with platform teams to ensure BioNeMo inference products deploy cleanly across cloud and enterprise environments, including Kubernetes-based environments.
- Develop the developer experience across APIs, SDKs, containers, Helm charts, reference architectures, documentation, and examples.
- Engage directly with early customers and partners to understand workflows, validate product direction, and turn feedback into prioritized requirements.
- Establish product metrics for adoption, usability, performance, and operational quality; use data and customer insight to improve the product.
- Drive cross-functional product launches with marketing, developer relations, sales, and solution architects.
- Help shape the product boundary between open-source biomolecular models, NVIDIA NIMs, and enterprise-ready deployment and support!
- 5+ years of industry experience, ideally including 3+ years in software engineering or a deeply technical role and 2+ years in product management.
- Master’s degree (or equivalent experience) in Electrical, Mechanical, Materials, or other related fields.
- Demonstrated experience defining and shipping technical products, platforms, APIs, SDKs, developer tools, or cloud services.
- Strong understanding of AI/ML inference systems and deployment tradeoffs, including model serving, GPU infrastructure, batching, throughput, latency, scaling, and cost-performance optimization.
- Familiarity with containers and cloud-native infrastructure, including Docker, Kubernetes, Helm, CI/CD, and public-cloud deployment concepts.
- Ability to work fluently with engineers on system architecture, performance bottlenecks, operational requirements, and roadmap tradeoffs.
- Strong customer discovery, product requirement definition, prioritization, and roadmap-management skills.
- Excellent written and verbal communication skills for both technical and non-technical audiences.
- A collaborative, self-starting approach and the ability to influence across highly matrixed teams.
- Experience launching and scaling AI inference, model-serving, MLOps, or GPU-accelerated infrastructure products.
- Hands‑on expertise with modern AI inference and orchestration platforms such as NVIDIA NIM, Triton, TensorRT‑LLM, vLLM, Ray, KServe, or Kubeflow.
- Proven success optimizing production inference workloads, including performance, scalability, observability, and multi‑tenant serving.
- Experience building developer platforms or open‑source products with strong ecosystem adoption and integrations.
- Ability to partner with research teams to productize frontier AI models, including familiarity with biomolecular and scientific foundation models.
NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward‑thinking and hardworking people in the world working for us. If you're creative and autonomous, we want to hear from you!
The base salary range is 148,000 USD - 224,250 USD.
You will also be eligible for equity and benefits.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer.
As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
NVIDIA pioneered accelerated computing.
Today, our AI infrastructure powers global intelligence, transforming every industry.
Learn more about NVIDIA.
About Nvidia
Sourced by ZipRecruiter
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology--and amazing people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what's never been done before takes vision, innovation, and the world's best talent.
Industry
Computer and electronic product manufacturing
Company size
10,000+ Employees
Headquarters location
Santa Clara, CA, US