Machine Learning Engineer
Mesa, AZ · On-site
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
Mesa, AZ · On-site
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
Mesa, AZ · On-site
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
Chandler, AZ · On-site
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
Chandler, AZ · On-site
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
Tucson, AZ · On-site
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
Tucson, AZ · On-site
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
Chandler, AZ · On-site
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
Chandler, AZ · On-site
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
Tucson, AZ · On-site
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
Tucson, AZ · On-site
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
Flagstaff, AZ · On-site
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
Flagstaff, AZ · On-site
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
Glendale, AZ · On-site
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
Glendale, AZ · On-site
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
Phoenix, AZ · On-site
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
Phoenix, AZ · On-site
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
Mesa, AZ · On-site
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
Mesa, AZ · On-site
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
Scottsdale, AZ · On-site
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
Scottsdale, AZ · On-site
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
Scottsdale, AZ · On-site
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
Scottsdale, AZ · On-site
* Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch * Build and maintain the infrastructure around RL training: rollout collection ...
Chandler, AZ · Remote
$51.76 - $77.74/hr
At Poe, we use Machine Learning in various parts of the product - bot routing, agent flow, code editing, RAG, etc. Our team of Machine Learning Engineers have high impact by advancing the current ...
Chandler, AZ · Remote
$51.76 - $77.74/hr
At Poe, we use Machine Learning in various parts of the product - bot routing, agent flow, code editing, RAG, etc. Our team of Machine Learning Engineers have high impact by advancing the current ...
Glendale, AZ · Remote
$51.76 - $77.74/hr
At Poe, we use Machine Learning in various parts of the product - bot routing, agent flow, code editing, RAG, etc. Our team of Machine Learning Engineers have high impact by advancing the current ...
Glendale, AZ · Remote
$51.76 - $77.74/hr
At Poe, we use Machine Learning in various parts of the product - bot routing, agent flow, code editing, RAG, etc. Our team of Machine Learning Engineers have high impact by advancing the current ...
Mesa, AZ · On-site
$139K - $168K/yr
At Poe, we use Machine Learning in various parts of the product - bot routing, agent flow, code editing, RAG, etc. Our team of Machine Learning Engineers have high impact by advancing the current ...
Mesa, AZ · On-site
$139K - $168K/yr
At Poe, we use Machine Learning in various parts of the product - bot routing, agent flow, code editing, RAG, etc. Our team of Machine Learning Engineers have high impact by advancing the current ...
Glendale, AZ · On-site
$139K - $168K/yr
At Poe, we use Machine Learning in various parts of the product - bot routing, agent flow, code editing, RAG, etc. Our team of Machine Learning Engineers have high impact by advancing the current ...
Glendale, AZ · On-site
$139K - $168K/yr
At Poe, we use Machine Learning in various parts of the product - bot routing, agent flow, code editing, RAG, etc. Our team of Machine Learning Engineers have high impact by advancing the current ...
Scottsdale, AZ · On-site
$139K - $168K/yr
At Poe, we use Machine Learning in various parts of the product - bot routing, agent flow, code editing, RAG, etc. Our team of Machine Learning Engineers have high impact by advancing the current ...
Scottsdale, AZ · On-site
$139K - $168K/yr
At Poe, we use Machine Learning in various parts of the product - bot routing, agent flow, code editing, RAG, etc. Our team of Machine Learning Engineers have high impact by advancing the current ...
Mesa, AZ · Remote
$51.76 - $77.74/hr
At Poe, we use Machine Learning in various parts of the product - bot routing, agent flow, code editing, RAG, etc. Our team of Machine Learning Engineers have high impact by advancing the current ...
Mesa, AZ · Remote
$51.76 - $77.74/hr
At Poe, we use Machine Learning in various parts of the product - bot routing, agent flow, code editing, RAG, etc. Our team of Machine Learning Engineers have high impact by advancing the current ...
Chandler, AZ · On-site
$139K - $168K/yr
At Poe, we use Machine Learning in various parts of the product - bot routing, agent flow, code editing, RAG, etc. Our team of Machine Learning Engineers have high impact by advancing the current ...
Chandler, AZ · On-site
$139K - $168K/yr
At Poe, we use Machine Learning in various parts of the product - bot routing, agent flow, code editing, RAG, etc. Our team of Machine Learning Engineers have high impact by advancing the current ...
Scottsdale, AZ · Remote
$51.76 - $77.74/hr
At Poe, we use Machine Learning in various parts of the product - bot routing, agent flow, code editing, RAG, etc. Our team of Machine Learning Engineers have high impact by advancing the current ...
Scottsdale, AZ · Remote
$51.76 - $77.74/hr
At Poe, we use Machine Learning in various parts of the product - bot routing, agent flow, code editing, RAG, etc. Our team of Machine Learning Engineers have high impact by advancing the current ...
The most popular types of Machine Learning jobs in Arizona are:
For Weekend Machine Learning jobs in Arizona, the most frequently searched job titles are:
The top searched job categories for Weekend Machine Learning jobs in Arizona are:
Cities in Arizona with the most Weekend Machine Learning job openings:
Full-time
Re-posted 22 days ago
Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch
Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment orchestration
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues
Build evaluation harnesses and benchmark infrastructure, with held-out sets and contamination controls, so results are trustworthy
Read eval signal and training curves to determine whether a change actually helped, and feed findings back to the research and environment teams
Integrate RL environments into the training stack, working with environment authors on interfaces, reward plumbing, and agent loop mechanics
Implement methods from recent ML papers quickly and turn them into production-grade systems