Apply reinforcement learning and model fine-tuning (e.g., instruction tuning, RLHF/RLAIF ... Collaborate with engineering, product management, and business teams to define requirements and ...
New
Apply reinforcement learning and model fine-tuning (e.g., instruction tuning, RLHF/RLAIF ... Collaborate with engineering, product management, and business teams to define requirements and ...
New
Apply reinforcement learning and model fine-tuning (e.g., instruction tuning, RLHF/RLAIF ... Collaborate with engineering, product management, and business teams to define requirements and ...
New
New York, NY · Remote
$104K - $138K/yr
Key Responsibilities Technical Leadership & Team Management * Lead, mentor, and grow a high ... human feedback (RLHF) or synthetic data augmentation. * Establish model-drift detection and ...
Quick apply
New York, NY · Remote
$104K - $138K/yr
Key Responsibilities Technical Leadership & Team Management * Lead, mentor, and grow a high ... human feedback (RLHF) or synthetic data augmentation. * Establish model-drift detection and ...
New York, NY · On-site
$184K - $324K/yr
... to manage their own context in long-horizon tasks. This is applied research with direct product ... RLHF, GRPO, PPO, RLVR, reward modeling, RL scaling laws Code generation and coding agents ...
New York, NY · On-site
$184K - $324K/yr
... to manage their own context in long-horizon tasks. This is applied research with direct product ... RLHF, GRPO, PPO, RLVR, reward modeling, RL scaling laws Code generation and coding agents ...
New York, NY · On-site
$165K - $300K/yr
Techniques Engineering Location NY New York United States Business Investment Management Function ... Experience in agent evaluation frameworks, RLHF, RLVR, policy optimization, or synthetic data ...
New York, NY · On-site
$165K - $300K/yr
Techniques Engineering Location NY New York United States Business Investment Management Function ... Experience in agent evaluation frameworks, RLHF, RLVR, policy optimization, or synthetic data ...
New York, NY · On-site
$115K - $200K/yr
... management. * Provide technical leadership through design reviews, mentorship, and cross-team ... Experience with reinforcement learning, fine-tuning, or preference-based optimization (e.g., RLHF)
New York, NY · On-site
$115K - $200K/yr
... management. * Provide technical leadership through design reviews, mentorship, and cross-team ... Experience with reinforcement learning, fine-tuning, or preference-based optimization (e.g., RLHF)
New York, NY · Remote
$115K - $200K/yr
... management. * Provide technical leadership through design reviews, mentorship, and cross-team ... Experience with reinforcement learning, fine-tuning, or preference-based optimization (e.g., RLHF)
Quick apply
New York, NY · Remote
$115K - $200K/yr
... management. * Provide technical leadership through design reviews, mentorship, and cross-team ... Experience with reinforcement learning, fine-tuning, or preference-based optimization (e.g., RLHF)
New York, NY · On-site
$150K - $215K/yr
... state management (short and long-term memory). * Deep familiarity designing and implementing ... Exposure to dataset curation and post-training techniques (SFT, DPO, RLHF) on open-weight models ...
New York, NY · On-site
$150K - $215K/yr
... state management (short and long-term memory). * Deep familiarity designing and implementing ... Exposure to dataset curation and post-training techniques (SFT, DPO, RLHF) on open-weight models ...
... managed platforms: Slurm, AWS PCS (Parallel Computing Service), SageMaker HyperPod - Experience with fine-tuning techniques: LoRA, QLoRA, RLHF, DPO, knowledge distillation, Quantization, KV ...
... managed platforms: Slurm, AWS PCS (Parallel Computing Service), SageMaker HyperPod - Experience with fine-tuning techniques: LoRA, QLoRA, RLHF, DPO, knowledge distillation, Quantization, KV ...
New York, NY · On-site
$100K - $136K/yr
... tuning, RLHF/RLAIF, preference optimization) to adapt large language models to our domain and ... We partner closely with strong engineers, product managers, and sales leaders who bring ads ...
New York, NY · On-site
$100K - $136K/yr
... tuning, RLHF/RLAIF, preference optimization) to adapt large language models to our domain and ... We partner closely with strong engineers, product managers, and sales leaders who bring ads ...
Livingston, NJ · On-site
$233K - $341K/yr
Build the infrastructure required for sophisticated Reinforcement Learning (RL) and RLHF pipelines ... managing large-scale infrastructure at a top-tier research lab or an AI-native cloud provider.
Livingston, NJ · On-site
$233K - $341K/yr
Build the infrastructure required for sophisticated Reinforcement Learning (RL) and RLHF pipelines ... managing large-scale infrastructure at a top-tier research lab or an AI-native cloud provider.
... self-managed on EKS, SageMaker endpoints, Sagemaker Hyperpod Serving), batching strategies ... RLHF/DPO, knowledge distillation, and efficient serving techniques (vLLM, TensorRT-LLM, Triton ...
... self-managed on EKS, SageMaker endpoints, Sagemaker Hyperpod Serving), batching strategies ... RLHF/DPO, knowledge distillation, and efficient serving techniques (vLLM, TensorRT-LLM, Triton ...
... self-managed on EKS, SageMaker endpoints, Sagemaker Hyperpod Serving), batching strategies ... RLHF/DPO, knowledge distillation, and efficient serving techniques (vLLM, TensorRT-LLM, Triton ...
... self-managed on EKS, SageMaker endpoints, Sagemaker Hyperpod Serving), batching strategies ... RLHF/DPO, knowledge distillation, and efficient serving techniques (vLLM, TensorRT-LLM, Triton ...
New York, NY · On-site
$264K - $369K/yr
We believe consumers should be able to understand and manage their financial life through ... RLHF) can power an insights flywheel; pioneer the architecture for customer-specific long-term ...
New York, NY · On-site
$264K - $369K/yr
We believe consumers should be able to understand and manage their financial life through ... RLHF) can power an insights flywheel; pioneer the architecture for customer-specific long-term ...
Manhattan, NY · On-site
... managers without being a direct people leader. You will be expected to be an external leader ... RLHF. * An engineering mindset as shown by a track record of delivering models at scale both in ...
Manhattan, NY · On-site
... managers without being a direct people leader. You will be expected to be an external leader ... RLHF. * An engineering mindset as shown by a track record of delivering models at scale both in ...
This is a people manager role that will lead teams to drive strategic direction through ... RLHF. * An engineering mindset as shown by a track record of delivering models at scale both in ...
This is a people manager role that will lead teams to drive strategic direction through ... RLHF. * An engineering mindset as shown by a track record of delivering models at scale both in ...
This is a people manager role that will lead teams to drive strategic direction through ... RLHF. * An engineering mindset as shown by a track record of delivering models at scale both in ...
This is a people manager role that will lead teams to drive strategic direction through ... RLHF. * An engineering mindset as shown by a track record of delivering models at scale both in ...
New York, NY · On-site
... managers without being a direct people leader. You will be expected to be an external leader ... RLHF. * An engineering mindset as shown by a track record of delivering models at scale both in ...
New York, NY · On-site
... managers without being a direct people leader. You will be expected to be an external leader ... RLHF. * An engineering mindset as shown by a track record of delivering models at scale both in ...
... managers without being a direct people leader. You will be expected to be an external leader ... RLHF. * An engineering mindset as shown by a track record of delivering models at scale both in ...
... managers without being a direct people leader. You will be expected to be an external leader ... RLHF. * An engineering mindset as shown by a track record of delivering models at scale both in ...
New York, NY · On-site
... managers without being a direct people leader. You will be expected to be an external leader ... RLHF. * An engineering mindset as shown by a track record of delivering models at scale both in ...
New York, NY · On-site
... managers without being a direct people leader. You will be expected to be an external leader ... RLHF. * An engineering mindset as shown by a track record of delivering models at scale both in ...
New York, NY · On-site
... RLHF, and DPO • experience optimizing workflows that rely on rate-limited APIs • you're curious ... We also provide diversified asset management solutions and focused investment banking capabilities.
New York, NY · On-site
... RLHF, and DPO • experience optimizing workflows that rely on rate-limited APIs • you're curious ... We also provide diversified asset management solutions and focused investment banking capabilities.
$25.4K - $34K
9% of jobs
$34K - $42.6K
15% of jobs
$43.3K is the 25th percentile. Wages below this are outliers.
$42.6K - $51.2K
17% of jobs
The median wage is $54.1K / yr.
$51.2K - $59.8K
27% of jobs
$65.1K is the 75th percentile. Wages above this are outliers.
$59.8K - $68.4K
12% of jobs
$68.4K - $77K
8% of jobs
$77K - $85.6K
4% of jobs
$85.6K - $94.3K
3% of jobs
$94.3K - $102.9K
2% of jobs
$102.9K - $111.5K
2% of jobs
$111.5K - $120.1K
1% of jobs
$25.4K
$61.6K
$120.1K
The most popular types of Rlhf jobs in Edison, NJ are:
The top searched job categories for Manager Rlhf jobs in Edison, NJ are:

$100K - $136K/yr
Full-time
Posted 3 days ago
New
7.4
Based on 7,147 frontline employees who took The Breakroom Quiz
5th of 39 rated national retailers
The Ads Marketing Decision Science team builds intelligent, data-driven systems that transform advertiser experiences through precise personalization and automated optimization. We decode complex patterns in advertiser behavior, content effectiveness, and performance signals to power real-time, contextual marketing decisions at scale - moving Amazon Ads from rules-based relevancy to true AI-driven personalization. Our work spans four pillars: Advertiser DNA (behavioral fingerprinting to predict advertiser needs and growth opportunities), Content Intelligence (frameworks to evaluate, select, and generate marketing content aligned to advertiser context), Automated Decision Systems (ML-powered audience targeting and next-best-action recommendations), and Gen-AI Applications (contextual, natural interactions across marketing touchpoints).
As a Senior Applied Scientist on the team, you will be at the forefront of our Gen-AI applications, leading the science behind conversational and agentic experiences that help advertisers grow
This role demands a strong foundation in machine learning and in LLM/NLP - deep fundamentals that you apply to build robust, production-grade systems rather than treating models as black boxes. In particular, you will own the development of our chatbot capability - designing the agentic reasoning, retrieval, and evaluation systems that make these interactions accurate, helpful, and trustworthy. You will set the technical vision, innovate on behalf of our customers, and take solutions end-to-end from inception to production.
You will partner closely with engineering to deploy at scale and low latency, and with product and business teams to ensure the experience meets real advertiser needs.
Key job responsibilities
Lead the design and development of the chatbot/agentic AI capability for WeChat and other third-party channels, from concept through production.
Bring strong ML and LLM/NLP fundamentals to bear on system design - grounding architecture and modeling choices in a deep understanding of the underlying methods.
Architect and build agentic AI systems - planning, tool use, and multi-step reasoning - grounded in Retrieval-Augmented Generation (RAG) over Amazon Ads knowledge sources.
Apply reinforcement learning and model fine-tuning (e.g., instruction tuning, RLHF/RLAIF, preference optimization) to adapt large language models to our domain and channels.
Define and operationalize rigorous LLM evaluation: golden sets, faithfulness/groundedness, precision/recall, and human-in-the-loop evaluation mechanisms that reliably measure and improve quality.
Own applied engineering quality of the science stack - PyTorch modeling, well-designed APIs, and latency/cost optimization for real-time, production-grade interactions.
Collaborate with engineering, product management, and business teams to define requirements and ship measurable customer impact.
Drive continuous improvement through experimentation, iterative development, testing, and optimization.
Translate complex scientific challenges into clear, impactful solutions for business stakeholders.
Mentor and guide junior scientists, fostering a collaborative, high-performing team culture, and engage the broader scientific community through presentations, publications, and patents.
About the team
We are a team of Applied Scientists, Research Scientists, Data Scientists, and Business Intelligence Engineers with deep expertise in ML, NLP, Gen-AI, RL, and causal inference, from a diverse range of backgrounds. We partner closely with strong engineers, product managers, and sales leaders who bring ads-industry depth and experience building scalable modeling and software solutions.
Sourced by ZipRecruiter
Amazon.com, Inc., commonly known as Amazon, is an American multinational technology company. It was founded by Jeff Bezos in 1994 and initially started as an online marketplace for books. Since then, Amazon has expanded its operations and become one of the largest e-commerce companies in the world. Amazon's primary business is its online retail platform, where customers can purchase a vast array of products, including electronics, clothing, books, home goods, and much more. The company offers a convenient and user-friendly shopping experience, with features such as fast shipping, customer reviews, and personalized recommendations. In addition to its e-commerce platform, Amazon has diversified its business into various other areas. One of its notable ventures is Amazon Web Services (AWS), a comprehensive cloud computing platform that provides services such as storage, compute power, and database management to individuals and businesses. AWS has become a leader in the cloud computing industry, powering many websites and applications worldwide. Amazon has also developed its own consumer electronics, including the popular Amazon Kindle e-reader, Fire tablets, Fire TV streaming devices, and the Alexa-powered Echo smart speakers. The Alexa voice assistant, integrated into these devices, allows users to interact with their devices using voice commands, perform tasks, and access information. Furthermore, Amazon has expanded into media and entertainment. It operates Prime Video, a streaming service that offers a wide range of movies, TV shows, and original content. Amazon Music provides a platform for streaming and purchasing digital music, while Audible offers audiobooks and other audio content. The company's commitment to customer satisfaction and convenience is demonstrated by its membership program, Amazon Prime. Prime members receive various benefits, including free two-day shipping, access to streaming services, exclusive deals, and more.
It services, book publishers, retail, real estate, computer and electronic product manufacturing and software development
10,000+ Employees
Seattle, WA, US