Own the execution of key AI/LLM benchmarks, including setup, orchestration, result collection, and analysis. * Investigate and resolve issues when training jobs or benchmarks fail, hang, or ...
Own the execution of key AI/LLM benchmarks, including setup, orchestration, result collection, and analysis. * Investigate and resolve issues when training jobs or benchmarks fail, hang, or ...
Run key AI/LLM benchmarks -- setup, orchestration, result collection, and analysis. * Investigate and address problems when training jobs or benchmarks fail, hang, or perform below expectations.
Run key AI/LLM benchmarks -- setup, orchestration, result collection, and analysis. * Investigate and address problems when training jobs or benchmarks fail, hang, or perform below expectations.
Run key AI/LLM benchmarks - setup, orchestration, result collection, and analysis. * Investigate and address problems when training jobs or benchmarks fail, hang, or perform below expectations.
Run key AI/LLM benchmarks - setup, orchestration, result collection, and analysis. * Investigate and address problems when training jobs or benchmarks fail, hang, or perform below expectations.
Experience with LLM training, fine-tuning (e.g., LoRa), and/or developing custom evaluation benchmarks. * Familiarity with advanced AI architectures and concepts such as agentic systems (e.g., ReAct ...
Experience with LLM training, fine-tuning (e.g., LoRa), and/or developing custom evaluation benchmarks. * Familiarity with advanced AI architectures and concepts such as agentic systems (e.g., ReAct ...
Experience with LLM training, fine-tuning (e.g., LoRa), and/or developing custom evaluation benchmarks. * Familiarity with advanced AI architectures and concepts such as agentic systems (e.g., ReAct ...
Experience with LLM training, fine-tuning (e.g., LoRa), and/or developing custom evaluation benchmarks. * Familiarity with advanced AI architectures and concepts such as agentic systems (e.g., ReAct ...
Experience with LLM training, fine-tuning (e.g., LoRa), and/or developing custom evaluation benchmarks. * Familiarity with advanced AI architectures and concepts such as agentic systems (e.g., ReAct ...
Quick apply
Experience with LLM training, fine-tuning (e.g., LoRa), and/or developing custom evaluation benchmarks. * Familiarity with advanced AI architectures and concepts such as agentic systems (e.g., ReAct ...
Senior Software Engineer Applied AI
Raleigh, NC · Remote
$119K - $157K/yr
Agentic tooling that gives AI systems safe, structured access to infrastructure Traditional (non-LLM) machine learning * End-to-end ML pipelines: feature engineering, model training, and scheduled ...
Senior Software Engineer Applied AI
Raleigh, NC · Remote
$119K - $157K/yr
Agentic tooling that gives AI systems safe, structured access to infrastructure Traditional (non-LLM) machine learning * End-to-end ML pipelines: feature engineering, model training, and scheduled ...
AI Engineer Consultant
Raleigh, NC · Hybrid
This role is hands-on and delivery-oriented: you will ship production pipelines and services that support model training, real-time inference, and LLM applications using Claude-, GPT/Codex-, and ...
AI Engineer Consultant
Raleigh, NC · Hybrid
This role is hands-on and delivery-oriented: you will ship production pipelines and services that support model training, real-time inference, and LLM applications using Claude-, GPT/Codex-, and ...
Establish governance practices for the context that drives AI/LLM outputs, including: * Curating ... Develop training and enablement to raise data and context literacy across the organization. * Lead ...
Establish governance practices for the context that drives AI/LLM outputs, including: * Curating ... Develop training and enablement to raise data and context literacy across the organization. * Lead ...
Staff AI Engineer
Raleigh, NC · On-site +1
... skill sets; experience and training; licensure and certifications; and other business and ... Advanced knowledge of LLM fine-tuning, alignment techniques, and evaluation methodologies including ...
Staff AI Engineer
Raleigh, NC · On-site +1
... skill sets; experience and training; licensure and certifications; and other business and ... Advanced knowledge of LLM fine-tuning, alignment techniques, and evaluation methodologies including ...
Staff AI Engineer
Raleigh, NC · On-site
... skill sets; experience and training; licensure and certifications; and other business and ... Advanced knowledge of LLM fine-tuning, alignment techniques, and evaluation methodologies including ...
Staff AI Engineer
Raleigh, NC · On-site
... skill sets; experience and training; licensure and certifications; and other business and ... Advanced knowledge of LLM fine-tuning, alignment techniques, and evaluation methodologies including ...
Staff AI Engineer
Raleigh, NC · On-site
... skill sets; experience and training; licensure and certifications; and other business and ... Advanced knowledge of LLM fine-tuning, alignment techniques, and evaluation methodologies including ...
Staff AI Engineer
Raleigh, NC · On-site
... skill sets; experience and training; licensure and certifications; and other business and ... Advanced knowledge of LLM fine-tuning, alignment techniques, and evaluation methodologies including ...
Senior Software Engineer Applied AI
Raleigh, NC · On-site
$150 - $210/hr
This is one seat that spans four disciplines that rarely come together: real‑time systems, LLM ... End‑to‑end ML pipelines: feature engineering, model training, and scheduled inference
Senior Software Engineer Applied AI
Raleigh, NC · On-site
$150 - $210/hr
This is one seat that spans four disciplines that rarely come together: real‑time systems, LLM ... End‑to‑end ML pipelines: feature engineering, model training, and scheduled inference
Dir Software Engineering
Raleigh, NC · On-site
$136K - $252K/yr
Practical understanding of LLM-based application development, including retrieval, evaluation ... Ensure staff receive clear goals, frequent feedback, coaching, training, and resources needed to ...
Dir Software Engineering
Raleigh, NC · On-site
$136K - $252K/yr
Practical understanding of LLM-based application development, including retrieval, evaluation ... Ensure staff receive clear goals, frequent feedback, coaching, training, and resources needed to ...
Dir Software Engineering
$136K - $252K/yr
Practical understanding of LLM-based application development, including retrieval, evaluation ... Ensure staff receive clear goals, frequent feedback, coaching, training, and resources needed to ...
Dir Software Engineering
$136K - $252K/yr
Practical understanding of LLM-based application development, including retrieval, evaluation ... Ensure staff receive clear goals, frequent feedback, coaching, training, and resources needed to ...
Dir Software Engineering
$136K - $252K/yr
Practical understanding of LLM-based application development, including retrieval, evaluation ... Ensure staff receive clear goals, frequent feedback, coaching, training, and resources needed to ...
Dir Software Engineering
$136K - $252K/yr
Practical understanding of LLM-based application development, including retrieval, evaluation ... Ensure staff receive clear goals, frequent feedback, coaching, training, and resources needed to ...
Dir Software Engineering
$136K - $252K/yr
Practical understanding of LLM-based application development, including retrieval, evaluation ... Ensure staff receive clear goals, frequent feedback, coaching, training, and resources needed to ...
Dir Software Engineering
$136K - $252K/yr
Practical understanding of LLM-based application development, including retrieval, evaluation ... Ensure staff receive clear goals, frequent feedback, coaching, training, and resources needed to ...
Principal AI Architect
Raleigh, NC · On-site
$160K - $240K/yr
Define and own enterprise AI and GenAI reference architectures, including LLM platforms, RAG ... Define enterprise patterns for training, fine-tuning, deployment, and observability of ML and GenAI ...
Principal AI Architect
Raleigh, NC · On-site
$160K - $240K/yr
Define and own enterprise AI and GenAI reference architectures, including LLM platforms, RAG ... Define enterprise patterns for training, fine-tuning, deployment, and observability of ML and GenAI ...
Principal AI Architect
Raleigh, NC · On-site
$160K - $240K/yr
Define and own enterprise AI and GenAI reference architectures, including LLM platforms, RAG ... Define enterprise patterns for training, fine-tuning, deployment, and observability of ML and GenAI ...
Principal AI Architect
Raleigh, NC · On-site
$160K - $240K/yr
Define and own enterprise AI and GenAI reference architectures, including LLM platforms, RAG ... Define enterprise patterns for training, fine-tuning, deployment, and observability of ML and GenAI ...
Dir Software Engineering
Raleigh, NC · On-site
$136K - $252K/yr
Practical understanding of LLM-based application development, including retrieval, evaluation ... Ensure staff receive clear goals, frequent feedback, coaching, training, and resources needed to ...
Dir Software Engineering
Raleigh, NC · On-site
$136K - $252K/yr
Practical understanding of LLM-based application development, including retrieval, evaluation ... Ensure staff receive clear goals, frequent feedback, coaching, training, and resources needed to ...
Llm Trainer information
See Raleigh, NC salary details
$18.88 is the 25th percentile. Wages below this are outliers.
$15.19 - $21.96
46% of jobs
The median wage is $23.05 / hr.
$21.96 - $28.74
26% of jobs
$28.74 - $35.52
0% of jobs
$35.52 - $42.29
0% of jobs
$47.37 is the 75th percentile. Wages above this are outliers.
$42.29 - $49.07
4% of jobs
$49.07 - $55.85
13% of jobs
$55.85 - $62.62
0% of jobs
$62.62 - $69.40
0% of jobs
$69.40 - $76.17
3% of jobs
$76.17 - $82.95
4% of jobs
$82.95 - $89.73
4% of jobs
$15
$35
$89
How much do llm trainer jobs pay per hour?
What does an LLM trainer do?
LLM Trainers are responsible for designing and refining training datasets, developing prompts, evaluating model outputs, and working closely with engineers and data scientists to optimize large language models. Common challenges include maintaining data quality, mitigating model biases, and staying up-to-date with rapidly evolving AI research and best practices. You’ll often collaborate with cross-functional teams, communicate findings clearly, and adapt to new tools or methodologies. This dynamic environment offers opportunities for innovation and skill development, making it an excellent fit for those passionate about advancing AI technology.
What skills and qualifications are needed to be an LLM trainer?
To thrive as an LLM Trainer, you need a deep understanding of natural language processing (NLP), machine learning principles, and data annotation techniques, often supported by a background in computer science or related fields. Familiarity with tools like Python, PyTorch or TensorFlow, data labeling platforms, and version control systems is essential, along with knowledge of prompt engineering and model fine-tuning. Strong analytical thinking, attention to detail, and collaborative communication skills are crucial soft skills for working with cross-functional AI teams. These competencies are important for developing high-quality language models that meet user needs and industry standards.
What is an LLM trainer?
An LLM Trainer is responsible for training and fine-tuning large language models (LLMs) to improve their accuracy, efficiency, and relevance for specific applications. This role involves curating and preprocessing training data, designing training methodologies, and evaluating model performance. LLM Trainers work closely with data scientists, engineers, and researchers to optimize models for tasks such as natural language understanding, text generation, and conversational AI. They also ensure ethical AI practices by mitigating biases and refining model outputs.

Full-time
Re-posted 14 days ago
Nvidia rating
9.6
Based on 17 frontline employees who took The Breakroom Quiz
8th of 242 rated software companies
Job description
We are seeking an ambitious Senior Solutions Architect - AI Factory Deployment to join our NVIDIA Infrastructure Specialists team in Santa Clara! This role is uniquely positioned to develop, deploy, and validate AI factories end to end. You will focus on running and debugging AI/LLM workloads and benchmarks on Linux-based GPU clusters, using NCCL and collectives like AllReduce and AllToAll to improve performance and scalability.
As part of our world-class team, you will bring to bear observability and automation to improve benchmarks and validation. You will serve as the expert when workloads or benchmarks do not perform flawlessly. You will collaborate across NVIDIA to ensure AI factories are prepared for customers, validating hardware and software for modern AI deployments.
What You Will be Doing:
Set up, adjust, and verify AI factory environments across multi-GPU and multi-node Linux clusters.
Ensure configurations align with guidelines for NCCL, collectives, and distributed training frameworks.
Own the execution of key AI/LLM benchmarks, including setup, orchestration, result collection, and analysis.
Investigate and resolve issues when training jobs or benchmarks fail, hang, or underperform.
Build and improve observability for AI factories (metrics, logs, traces, dashboards) to understand workload behavior and system health.
Develop automation (Python, Shell) for running benchmarks, collecting results, and performing regression checks
Examine communication patterns and NCCL usage for AI/LLM workloads, concentrating on collectives such as AllReduce and AllToAll.
Recommend changes to job configuration, parallelism strategies, and cluster settings to improve throughput, latency, and scaling efficiency.
Work closely with hardware, software, networking, datacenter, and product teams to prepare AI factories for customer use.
Contribute to documentation, guidelines, and readiness collateral that support internal collaborators and customer-facing teams.
What We Need to See:
Bachelor's degree or equivalent experience in Computer Science, Mathematics, Engineering, Physics, or related field.
More than 6+ years of experience managing Linux-based systems in HPC, distributed systems, or extensive AI/ML settings.
Hands-on experience running AI/ML workloads on multi-GPU and/or multi-node clusters, with practical knowledge of NCCL.
Solid grasp of collective communication patterns, particularly AllReduce and AllToAll, and how they are applied in contemporary ML/LLM training.
Familiarity with LLM training and/or inference workflows using frameworks such as PyTorch or TensorFlow.
Proficiency with Python and Shell/Bash for scripting, automation, and tooling.
Experience with benchmarking (crafting, executing, and interpreting performance benchmarks).
Comfortable working with observability data (metrics, logs, dashboards) to troubleshoot and optimize complex distributed workloads.
Strong communication skills and the ability to work effectively with cross-functional teams.
Ways to Stand Out From the Crowd:
Experience with AI factory or large-scale AI infrastructure build, deployment, or operations.
Background in HPC performance engineering, SRE, or systems performance analysis for GPU-accelerated environments.
Familiarity with observability stacks (e.g., metrics/monitoring, logging, tracing systems) used for large distributed systems.
Experience building automation and CI-style pipelines for running and validating benchmarks at scale.
Demonstrated desire to use AI to solve practical problems, improve workflows, and guide data-driven decisions.
You will also be eligible for equity and benefits.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.About Nvidia
Sourced by ZipRecruiter
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology--and amazing people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what's never been done before takes vision, innovation, and the world's best talent.
Industry
Computer and electronic product manufacturing
Company size
10,000+ Employees
Headquarters location
Santa Clara, CA, US
Year founded
1993