1

Rte Release Train Engineer Jobs in Santa Rosa, CA

Scientific Data Engineer

Bodega Bay, CA · On-site

$135K - $163K/yr

... software releases, continuous integration, and testing * Design, implement and maintain high ... Train scientists and research software engineers in the use of the developed software products at ...

Test Technician III

Petaluma, CA · On-site

$23 - $35/hr

... released engineering drawings and assembly instructions using our work order system. Position ... May also train other technicians to build a product or follow an established assembly process * May ...

... released engineering drawings and assembly instructions using our work order system. Position ... May also train other technicians to build a product or follow an established assembly process * May ...

PCB Designer

Bodega Bay, CA · On-site

$123K - $218K/yr

Our Engineering PCB Services Organization takes pride in supporting board designs for Electrical ... Train PCB Designers in specialized areas of PCB design. Developing an Understanding of Apple ...

Rte Release Train Engineer information

See Santa Rosa, CA salary details

$32

$68

$100

How much do rte release train engineer jobs pay per hour?

As of Sep 2, 2026, the average hourly pay for rte release train engineer in Santa Rosa, CA is $68.35, according to ZipRecruiter salary data. Most workers in this role earn between $56.78 and $76.20 per hour, depending on experience, location, and employer.

What is a release train engineer?

Release Train Engineers (RTEs) are servant leaders and coaches for Agile Release Trains (ARTs) within the Scaled Agile Framework (SAFe). They facilitate ART processes, coordinate between teams, and ensure value delivery by removing impediments and fostering collaboration. RTEs help manage risks, drive continuous improvement, and support teams in delivering solutions that align with organizational goals. Their role is vital in orchestrating large-scale Agile initiatives across multiple teams.

How does a release train engineer facilitate collaboration across multiple Agile teams within an Agile Release Train?

A Release Train Engineer (RTE) plays a pivotal role in coordinating and aligning several Agile teams working within an Agile Release Train (ART). They facilitate program-level ceremonies such as PI Planning, Scrum of Scrums, and System Demos to ensure transparency and alignment on objectives and deliverables. The RTE acts as a servant leader, removing impediments, fostering open communication, and supporting cross-team problem-solving. This role requires strong organizational and interpersonal skills to manage dependencies, risks, and continuous improvement across teams, ensuring that the ART delivers value effectively.

What are the key skills and qualifications needed to thrive as a release train engineer, and why are they important?

To thrive as a Release Train Engineer, you need expertise in Agile methodologies, project management, and a strong understanding of the Scaled Agile Framework (SAFe), often supported by a SAFe Release Train Engineer certification. Familiarity with Agile project management tools such as Jira, Confluence, and other collaboration systems is typically required. Exceptional communication, facilitation, and conflict resolution skills help foster collaboration and drive alignment across multiple teams. These abilities are crucial for coordinating large-scale Agile initiatives, removing obstacles, and ensuring successful program delivery.

What is the difference between Rte Release Train Engineer vs Scrum Master?

AspectRte Release Train EngineerScrum Master
CertificationsSAFe Program Consultant (SPC), Agile certificationsCertified ScrumMaster (CSM), SAFe certifications
Work EnvironmentLarge-scale Agile/SAFe environments, multiple teamsSingle Scrum team, Agile projects
Employer & Industry UsageOrganizations adopting SAFe, enterprise-levelStartups, small teams, Agile organizations

The Rte Release Train Engineer and Scrum Master both facilitate Agile practices, but the Rte focuses on coordinating multiple teams within a SAFe framework at an enterprise level, while the Scrum Master supports a single team to ensure Agile processes are followed effectively.

What are popular job titles related to Rte Release Train Engineer jobs in Santa Rosa, CA?

For Rte Release Train Engineer jobs in Santa Rosa, CA, the most frequently searched job titles are:

What cities near Santa Rosa, CA are hiring for Rte Release Train Engineer jobs?

Cities near Santa Rosa, CA with the most Rte Release Train Engineer job openings:

Machine Learning Engineer, LLM Inference Optimization

GMI Cloud

Sonoma, CA • On-site

Other

Re-posted 5 days ago


Job description

About Us

GMI Cloud is a fast-growing AI infrastructure company backed by Headline VC and one of only seven cloud providers worldwide to earn NVIDIA’s prestigious Reference Platform Cloud Partner designation . We operate 8 of our own GPU clusters across the U.S. and Asia, delivering a full spectrum of services from GPU compute service to AI model inference API solutions. As an NVIDIA Reference Platform Cloud Partner, our infrastructure meets the highest standards for performance, security, and scalability in AI deployments. We empower AI startups and enterprises to “build AI without limits,” providing everything they need to prototype, train, and deploy AI models quickly and reliably.


GMI Cloud is building the leading inference optimization solution and the most advanced token platform in the global token market — and we are hiring world-class Machine Learning Engineers to make GMI the new industry benchmark for LLM serving performance, cost efficiency, and production reliability.


This role is for engineers who want to live at the frontier of LLM inference systems. You will drive the research, validation, and productionization of the most advanced inference optimization techniques, and turn them into real competitive advantage over top open-source baselines (vLLM, SGLang, and so on). Our charter is not just to adopt what's published — it is to define the recipes, ship the optimizations, and contribute back to the community that the rest of the industry follows.


You will focus on B200-first optimization, with support for H200 evolution, across core domains including quantization, speculative decoding, KV cache and memory management, prefill/decode disaggregation, and system-level inference optimization. You will work closely with platform and infrastructure teams to transform cutting-edge ideas into measurable gains in latency, throughput, cost efficiency, and production scalability.


Key Responsibilities

  • Drive frontier research and engineering in LLM inference optimization across one of the four focus tracks (Speculative Decoding, Quantization, PD Disaggregation, KV Cache & Memory) while contributing across the full optimization stack.
  • Develop next-generation optimization strategies for large-scale LLM serving across model execution, runtime systems, and production inference platforms — with B200 as the primary target and H200 as a continuing platform.
  • Advance state-of-the-art techniques in quantization (NVFP4 / MXFP4 / FP8, QAT), speculative decoding (EAGLE-3, MTP, DFlash, ModelOpt, SpecForge), KV cache & memory management (LMCache / HiCache / NV KVBM, paged attention, prefix-aware routing), and PD disaggregation (NVIDIA Dynamo, KV-aware router/planner, fault recovery).
  • Drive system-level optimization across scheduling, batching, routing, gateway orchestration, adapter serving, and end-to-end inference efficiency.
  • Build scalable optimization frameworks, performance methodologies, and benchmark infrastructure that allow GMI to stay ahead of the industry as models, hardware, and serving patterns evolve.
  • Productionize cutting-edge ideas into real customer workloads — measured by TTFT, ITL, throughput, goodput, tail latency, quality, and unit token cost.
  • Engage with and contribute to the open-source community (vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo / ModelOpt, FlashInfer, LMCache, etc.) — read upstream code, file issues, send PRs, and publish tech blogs and case studies.
  • Collaborate closely with platform, infrastructure, and product teams to make inference optimization a core technical advantage of GMI Cloud.


Required Skills

  • Strong hands-on experience with LLM inference systems and performance optimization on modern GPUs.
  • Solid understanding of inference metrics and tradeoffs, including TTFT, ITL, throughput, goodput, tail latency, GPU utilization, memory efficiency, and quality/cost tradeoffs.
  • Experience with one or more modern serving stacks such as SGLang, vLLM, TensorRT-LLM, NVIDIA Dynamo, or Triton.
  • Deep familiarity with GPU-based inference, model serving architecture, and production bottlenecks around compute, memory bandwidth, KV-cache behavior, and scheduling.
  • Demonstrable depth in at least one of the four focus areas: speculative decoding, quantization & precision, PD disaggregation, or KV cache & memory management.
  • Strong experimentation skills: able to design benchmarks, interpret results, debug regressions, and produce actionable conclusions rather than isolated microbenchmark wins.
  • Proficient with Claude Code at an advanced level — fluent with sub-agents, MCP servers, hooks, custom slash commands, and skills — with practical experience leveraging them for rapid iteration, profiling, observability, and performance debugging.
  • Clear communication — able to explain technical tradeoffs to engineers and cross-functional stakeholders, and willing to publish results externally.


Preferred Qualifications

  • 2+ years of hands-on experience in LLM inference optimization, ML systems optimization, or PhD degree in related areas.
  • Track record of large-scale model serving optimization (latency reduction, throughput improvement, memory efficiency, cost-performance tuning) in production.
  • Specific track depth in one or more of:
  • Speculative Decoding: EAGLE-3 / MTP / DFlash / Medusa / SpecForge / ModelOpt; experience training and shipping draft models for production.
  • Quantization & Precision: NVFP4 / MXFP4 / FP8 / INT4-AWQ / GPTQ; QAT pipelines on Blackwell or Hopper; rigorous accuracy benchmarking.
  • PD Disaggregation: NVIDIA Dynamo, KV-aware router/planner, large MoE serving (DeepSeek-V3/V4, Kimi, GLM, Minimax), fault recovery, autoscaling.
  • KV Cache & Memory: LMCache / HiCache / NV KVBM, paged attention internals, prefix-aware routing, long-context and agentic workloads.
  • Familiarity with FlashInfer, Blackwell MLA, FA4, TRT-LLM MLA, or NSA is a strong plus.
  • Open-source contributions to vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo / ModelOpt, FlashInfer, LMCache, or related projects.
  • Experience publishing technical blogs, case studies, or papers on inference optimization.