... Machine Learning Engineers to make GMI the new industry benchmark for LLM serving performance, cost efficiency, and production reliability. This role is for engineers who want to live at the frontier ...
... Machine Learning Engineers to make GMI the new industry benchmark for LLM serving performance, cost efficiency, and production reliability. This role is for engineers who want to live at the frontier ...
... Machine Learning Engineers to make GMI the new industry benchmark for LLM serving performance, cost efficiency, and production reliability. This role is for engineers who want to live at the frontier ...
... Machine Learning Engineers to make GMI the new industry benchmark for LLM serving performance, cost efficiency, and production reliability. This role is for engineers who want to live at the frontier ...
... Machine Learning Engineers to make GMI the new industry benchmark for LLM serving performance, cost efficiency, and production reliability. This role is for engineers who want to live at the frontier ...
... Machine Learning Engineers to make GMI the new industry benchmark for LLM serving performance, cost efficiency, and production reliability. This role is for engineers who want to live at the frontier ...
Machine Learning Engineer: LLM, VLM/VLA and reasoning models (San Jose)
San Jose, CA · On-site
$300K/yr
Machine Learning Engineer: LLM, VLM/VLA and reasoning models Tensor is an agentic AI company dedicated to building agentic products that empower individual consumers. Our flagship product, the Tensor ...
Machine Learning Engineer: LLM, VLM/VLA and reasoning models (San Jose)
San Jose, CA · On-site
$300K/yr
Machine Learning Engineer: LLM, VLM/VLA and reasoning models Tensor is an agentic AI company dedicated to building agentic products that empower individual consumers. Our flagship product, the Tensor ...
Machine Learning Engineer
Burlington, MA · Remote
$165K - $200K/yr
Design and implement AI agents, agentic workflows, and LLM-powered applications. * Deploy and ... At MatrixSpace, Machine Learning Engineering is where advanced AI research becomes real-world ...
Quick apply
Machine Learning Engineer
Burlington, MA · Remote
$165K - $200K/yr
Design and implement AI agents, agentic workflows, and LLM-powered applications. * Deploy and ... At MatrixSpace, Machine Learning Engineering is where advanced AI research becomes real-world ...
Machine Learning Engineer
Richmond, VA · On-site
Machine Learning Engineer Richmond, Virginia (5 Days Onsite) need local within commute About the ... LLM-based agents using frameworks such as LangChain (or equivalent) Develop scalable backend ...
Quick apply
Machine Learning Engineer
Richmond, VA · On-site
Machine Learning Engineer Richmond, Virginia (5 Days Onsite) need local within commute About the ... LLM-based agents using frameworks such as LangChain (or equivalent) Develop scalable backend ...
About the Role We are looking for a hands-on Machine Learning Engineer to drive the post-training ... Hands-on LLM post-training experience. You have personally run CPT, SFT, and RL training - with ...
About the Role We are looking for a hands-on Machine Learning Engineer to drive the post-training ... Hands-on LLM post-training experience. You have personally run CPT, SFT, and RL training - with ...
Machine Learning Engineer, LLM Post-Training
Mountain View, CA · On-site
$150 - $200/hr
About the Role We are looking for a hands‑on Machine Learning Engineer to drive the ... Hands‑on LLM post‑training experience. You have personally run CPT, SFT, and RL training ...
Machine Learning Engineer, LLM Post-Training
Mountain View, CA · On-site
$150 - $200/hr
About the Role We are looking for a hands‑on Machine Learning Engineer to drive the ... Hands‑on LLM post‑training experience. You have personally run CPT, SFT, and RL training ...
Machine Learning Engineer
Cupertino, CA · On-site
Job Summary : Syntricate Technologies is a company seeking a Machine Learning Engineer with ... using LLM's • Managing Data pipeline/Transforms • Experience in Data injection and Data ...
Machine Learning Engineer
Cupertino, CA · On-site
Job Summary : Syntricate Technologies is a company seeking a Machine Learning Engineer with ... using LLM's • Managing Data pipeline/Transforms • Experience in Data injection and Data ...
... machine learning, and smart connectivity. Our mission is to build strong foundation for LLM ... Strong Python programming and software engineering skills. * Ability to work effectively across ...
... machine learning, and smart connectivity. Our mission is to build strong foundation for LLM ... Strong Python programming and software engineering skills. * Ability to work effectively across ...
... machine learning, and smart connectivity. Our mission is to build strong foundation for LLM ... Strong Python programming and software engineering skills. * Ability to work effectively across ...
... machine learning, and smart connectivity. Our mission is to build strong foundation for LLM ... Strong Python programming and software engineering skills. * Ability to work effectively across ...
Machine Learning Engineer
Santa Clara, CA · On-site
Our team comprises a diverse range of backgrounds, including applied machine learning engineers with a focus on ML and LLM, and experienced distributed systems engineers. As such, we are seeking ...
Machine Learning Engineer
Santa Clara, CA · On-site
Our team comprises a diverse range of backgrounds, including applied machine learning engineers with a focus on ML and LLM, and experienced distributed systems engineers. As such, we are seeking ...
Machine Learning Engineer
Los Angeles, CA · On-site
ROLE SUMMARY The Machine Learning Engineer is a major contributor in driving our company ... Conceptual knowledge of LLM's. * Strong knowledge of statistics, hypothesis testing, and setting up ...
Machine Learning Engineer
Los Angeles, CA · On-site
ROLE SUMMARY The Machine Learning Engineer is a major contributor in driving our company ... Conceptual knowledge of LLM's. * Strong knowledge of statistics, hypothesis testing, and setting up ...
NY · On-site
$60 - $80/hr
... LLM zgodnie z najlepszymi praktykami MLOps / LLMOps, * rozwój i utrzymanie środowisk ... Engineer / Machine Learning Engineer lub w podobnej roli, * bardzo dobrze znasz Python i masz ...
Machine Learning Engineer, LLM Post-Training
Mountain View, CA · On-site
$150K - $230K/yr
... Machine Learning Engineer to drive the post-training of our large language models, with a strong ... Hands-on LLM post-training experience. You have personally run CPT, SFT, and RL training - with ...
Machine Learning Engineer, LLM Post-Training
Mountain View, CA · On-site
$150K - $230K/yr
... Machine Learning Engineer to drive the post-training of our large language models, with a strong ... Hands-on LLM post-training experience. You have personally run CPT, SFT, and RL training - with ...
... machine learning, and smart connectivity. Our mission is to build strong foundation for LLM ... Strong Python programming and software engineering skills. * Ability to work effectively across ...
... machine learning, and smart connectivity. Our mission is to build strong foundation for LLM ... Strong Python programming and software engineering skills. * Ability to work effectively across ...
Machine Learning Engineer
Pittsburgh, PA · On-site
As a Machine Learning Engineer in the Machine Intelligence Neural Design (MIND) team, you'll have ... Strong foundation in machine learning, and more specifically in LLM and multimodal foundation ...
Machine Learning Engineer
Pittsburgh, PA · On-site
As a Machine Learning Engineer in the Machine Intelligence Neural Design (MIND) team, you'll have ... Strong foundation in machine learning, and more specifically in LLM and multimodal foundation ...
Are you a passionate Machine Learning Engineer with a strong background in SageMaker, prompt engineering, and LLM (Large Language Model) model tuning? Do you thrive in a dynamic and innovative ...
Quick apply
Are you a passionate Machine Learning Engineer with a strong background in SageMaker, prompt engineering, and LLM (Large Language Model) model tuning? Do you thrive in a dynamic and innovative ...
Machine Learning Engineer, LLM Post-Training
Mountain View, CA · On-site
$150K - $230K/yr
... Machine Learning Engineer to drive the post-training of our large language models, with a strong ... Hands-on LLM post-training experience. You have personally run CPT, SFT, and RL training -- with ...
Quick apply
Machine Learning Engineer, LLM Post-Training
Mountain View, CA · On-site
$150K - $230K/yr
... Machine Learning Engineer to drive the post-training of our large language models, with a strong ... Hands-on LLM post-training experience. You have personally run CPT, SFT, and RL training -- with ...
... machine learning, and smart connectivity. Our mission is to build strong foundation for LLM ... Strong Python programming and software engineering skills. * Ability to work effectively across ...
... machine learning, and smart connectivity. Our mission is to build strong foundation for LLM ... Strong Python programming and software engineering skills. * Ability to work effectively across ...
Machine Learning Engineer Llm information
See salary details
$31.5K - $46.2K
1% of jobs
$46.2K - $61K
1% of jobs
$61K - $75.7K
5% of jobs
$75.7K - $90.4K
6% of jobs
$102.6K is the 25th percentile. Wages below this are outliers.
$90.4K - $105.1K
14% of jobs
$105.1K - $119.9K
14% of jobs
The median wage is $127.2K / yr.
$119.9K - $134.6K
18% of jobs
$134.6K - $149.3K
14% of jobs
$152.3K is the 75th percentile. Wages above this are outliers.
$149.3K - $164K
12% of jobs
$164K - $178.8K
11% of jobs
$178.8K - $193.5K
5% of jobs
$31.5K
$128.8K
$193.5K
How much do machine learning engineer llm jobs pay per year?
What is a machine learning engineer LLM?
What are common challenges machine learning engineers face when working with large language models (LLMs) in a production environment?
What are the key skills and qualifications needed to thrive as a machine learning engineer LLM?
What is the difference between Machine Learning Engineer Llm vs Data Scientist?
| Aspect | Machine Learning Engineer Llm | Data Scientist |
|---|---|---|
| Required Credentials | Bachelor's or Master's in CS, AI, or related; experience with ML frameworks | Bachelor's or Master's in CS, Statistics, or related; strong analytical skills |
| Work Environment | Develops, tests, and deploys ML models, often in AI-focused teams | Analyzes data, builds models, and provides insights for decision-making |
| Industry Usage | Used in AI product development, NLP, LLMs, and automation | Applied across finance, healthcare, marketing, and research |
While both roles require strong technical skills and knowledge of machine learning, Machine Learning Engineer Llm focuses on developing and deploying large language models, especially in AI applications. Data Scientists analyze data and build models for insights. The roles often overlap but differ mainly in their focus on deployment versus analysis.
What are popular job titles related to Machine Learning Engineer Llm jobs?
For Machine Learning Engineer Llm jobs, the most frequently searched job titles are:

Machine Learning Engineer, LLM Inference Optimization
San Francisco, CA • On-site
Other
Re-posted 11 days ago
Job description
About Us
GMI Cloud is a fast-growing AI infrastructure company backed by Headline VC and one of only seven cloud providers worldwide to earn NVIDIA’s prestigious Reference Platform Cloud Partner designation . We operate 8 of our own GPU clusters across the U.S. and Asia, delivering a full spectrum of services from GPU compute service to AI model inference API solutions. As an NVIDIA Reference Platform Cloud Partner, our infrastructure meets the highest standards for performance, security, and scalability in AI deployments. We empower AI startups and enterprises to “build AI without limits,” providing everything they need to prototype, train, and deploy AI models quickly and reliably.
GMI Cloud is building the leading inference optimization solution and the most advanced token platform in the global token market — and we are hiring world-class Machine Learning Engineers to make GMI the new industry benchmark for LLM serving performance, cost efficiency, and production reliability.
This role is for engineers who want to live at the frontier of LLM inference systems. You will drive the research, validation, and productionization of the most advanced inference optimization techniques, and turn them into real competitive advantage over top open-source baselines (vLLM, SGLang, and so on). Our charter is not just to adopt what's published — it is to define the recipes, ship the optimizations, and contribute back to the community that the rest of the industry follows.
You will focus on B200-first optimization, with support for H200 evolution, across core domains including quantization, speculative decoding, KV cache and memory management, prefill/decode disaggregation, and system-level inference optimization. You will work closely with platform and infrastructure teams to transform cutting-edge ideas into measurable gains in latency, throughput, cost efficiency, and production scalability.
Key Responsibilities
- Drive frontier research and engineering in LLM inference optimization across one of the four focus tracks (Speculative Decoding, Quantization, PD Disaggregation, KV Cache & Memory) while contributing across the full optimization stack.
- Develop next-generation optimization strategies for large-scale LLM serving across model execution, runtime systems, and production inference platforms — with B200 as the primary target and H200 as a continuing platform.
- Advance state-of-the-art techniques in quantization (NVFP4 / MXFP4 / FP8, QAT), speculative decoding (EAGLE-3, MTP, DFlash, ModelOpt, SpecForge), KV cache & memory management (LMCache / HiCache / NV KVBM, paged attention, prefix-aware routing), and PD disaggregation (NVIDIA Dynamo, KV-aware router/planner, fault recovery).
- Drive system-level optimization across scheduling, batching, routing, gateway orchestration, adapter serving, and end-to-end inference efficiency.
- Build scalable optimization frameworks, performance methodologies, and benchmark infrastructure that allow GMI to stay ahead of the industry as models, hardware, and serving patterns evolve.
- Productionize cutting-edge ideas into real customer workloads — measured by TTFT, ITL, throughput, goodput, tail latency, quality, and unit token cost.
- Engage with and contribute to the open-source community (vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo / ModelOpt, FlashInfer, LMCache, etc.) — read upstream code, file issues, send PRs, and publish tech blogs and case studies.
- Collaborate closely with platform, infrastructure, and product teams to make inference optimization a core technical advantage of GMI Cloud.
Required Skills
- Strong hands-on experience with LLM inference systems and performance optimization on modern GPUs.
- Solid understanding of inference metrics and tradeoffs, including TTFT, ITL, throughput, goodput, tail latency, GPU utilization, memory efficiency, and quality/cost tradeoffs.
- Experience with one or more modern serving stacks such as SGLang, vLLM, TensorRT-LLM, NVIDIA Dynamo, or Triton.
- Deep familiarity with GPU-based inference, model serving architecture, and production bottlenecks around compute, memory bandwidth, KV-cache behavior, and scheduling.
- Demonstrable depth in at least one of the four focus areas: speculative decoding, quantization & precision, PD disaggregation, or KV cache & memory management.
- Strong experimentation skills: able to design benchmarks, interpret results, debug regressions, and produce actionable conclusions rather than isolated microbenchmark wins.
- Proficient with Claude Code at an advanced level — fluent with sub-agents, MCP servers, hooks, custom slash commands, and skills — with practical experience leveraging them for rapid iteration, profiling, observability, and performance debugging.
- Clear communication — able to explain technical tradeoffs to engineers and cross-functional stakeholders, and willing to publish results externally.
Preferred Qualifications
- 2+ years of hands-on experience in LLM inference optimization, ML systems optimization, or PhD degree in related areas.
- Track record of large-scale model serving optimization (latency reduction, throughput improvement, memory efficiency, cost-performance tuning) in production.
- Specific track depth in one or more of:
- Speculative Decoding: EAGLE-3 / MTP / DFlash / Medusa / SpecForge / ModelOpt; experience training and shipping draft models for production.
- Quantization & Precision: NVFP4 / MXFP4 / FP8 / INT4-AWQ / GPTQ; QAT pipelines on Blackwell or Hopper; rigorous accuracy benchmarking.
- PD Disaggregation: NVIDIA Dynamo, KV-aware router/planner, large MoE serving (DeepSeek-V3/V4, Kimi, GLM, Minimax), fault recovery, autoscaling.
- KV Cache & Memory: LMCache / HiCache / NV KVBM, paged attention internals, prefix-aware routing, long-context and agentic workloads.
- Familiarity with FlashInfer, Blackwell MLA, FA4, TRT-LLM MLA, or NSA is a strong plus.
- Open-source contributions to vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo / ModelOpt, FlashInfer, LMCache, or related projects.
- Experience publishing technical blogs, case studies, or papers on inference optimization.