1

Overnight Network Observability Jobs (NOW HIRING)

... network. You'll work at the intersection of large-scale distributed systems and cutting-edge ML ... observability for training workloads and model performance - Drive best practices for GPU fleet ...

Sr. Software Development Engineer, MLOPs

Bellevue, WA · On-site

$138K - $182K/yr

... network. You'll work at the intersection of large-scale distributed systems and cutting-edge ML ... observability for training workloads and model performance - Drive best practices for GPU fleet ...

Establish observability and monitoring standards across infrastructure, applications, and AI-driven ... Knowledge of networking technologies including WAN, VPN, WiFi, firewalls, routers, and Layer 2/3 ...

next page

Showing results 1-20

Overnight Network Observability information

What are the most commonly searched types of Network Observability jobs? The most popular types of Network Observability jobs are:
Infographic showing various Overnight Network Observability job openings in the United States as of June 2026, with employment types broken down into 1% Locum Tenens, 1% As Needed, 83% Full Time, and 15% Part Time. Highlights an 77% Physical, 7% Hybrid, and 16% Remote job distribution.
Sr. Software Development Engineer, MLOPs

Sr. Software Development Engineer, MLOPs

Amazon

Bellevue, WA

$138K - $182K/yr

Full-time

Posted 11 days ago


Amazon rating

7.4

Company rating: 7.4 out of 10

Based on 6,870 frontline employees who took The Breakroom Quiz

6th of 39 rated national retailers


Job description

We are looking for a Senior Software Development Engineer with deep expertise in machine learning operations to join the Data & Intelligence Foundation (DIF) team within Amazon. You will design, build, and operate the ML training infrastructure that enables robot learning at scale - from distributed GPU training pipelines to experiment tracking, data management, and model deployment.
Our team is building the foundational ML platform that powers autonomous robotics across Amazon's fulfillment network. You'll work at the intersection of large-scale distributed systems and cutting-edge ML research, turning novel vision-language-action models into production training workflows.
Key job responsibilities
- Design and implement scalable ML training infrastructure on Kubernetes (EKS) with GPU scheduling and fault-tolerant distributed training
- Build and maintain CI/CD pipelines for ML models - from data ingestion through training, evaluation, and deployment
- Develop tooling for experiment tracking, hyperparameter optimization, and reproducibility
- Architect data pipelines that handle large-scale robotics datasets (telemetry, sensor recordings, demonstrations)
- Collaborate with research scientists to operationalize novel ML models into production
- Establish monitoring, alerting, and observability for training workloads and model performance
- Drive best practices for GPU fleet management, cost optimization, and capacity planning
A day in the life
You'll spend your mornings reviewing training job health across our GPU cluster, debugging a distributed training run that hit a node failure overnight, and shipping a fix to our checkpoint recovery system

After lunch, you'll pair with a research scientist to optimize their new imitation learning model for multi-node training, then architect a new data pipeline to ingest demonstration recordings from robot workcells. You'll close the day reviewing a PR from a teammate and planning the next iteration of our experiment tracking platform.
About the team
The Data & Intelligence Foundation (DIF) team builds the ML infrastructure platform for Amazon's industrial robotics. We enable scientists to train, evaluate, and deploy models that power autonomous robots in fulfillment centers worldwide

We're a small, high-impact team where every engineer shapes the architecture and directly accelerates robot intelligence. We value pragmatic engineering, deep technical ownership, and close collaboration with research.


What Amazon employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom


Amazon logo

About Amazon

Sourced by ZipRecruiter

Amazon.com, Inc., commonly known as Amazon, is an American multinational technology company. It was founded by Jeff Bezos in 1994 and initially started as an online marketplace for books. Since then, Amazon has expanded its operations and become one of the largest e-commerce companies in the world. Amazon's primary business is its online retail platform, where customers can purchase a vast array of products, including electronics, clothing, books, home goods, and much more. The company offers a convenient and user-friendly shopping experience, with features such as fast shipping, customer reviews, and personalized recommendations. In addition to its e-commerce platform, Amazon has diversified its business into various other areas. One of its notable ventures is Amazon Web Services (AWS), a comprehensive cloud computing platform that provides services such as storage, compute power, and database management to individuals and businesses. AWS has become a leader in the cloud computing industry, powering many websites and applications worldwide. Amazon has also developed its own consumer electronics, including the popular Amazon Kindle e-reader, Fire tablets, Fire TV streaming devices, and the Alexa-powered Echo smart speakers. The Alexa voice assistant, integrated into these devices, allows users to interact with their devices using voice commands, perform tasks, and access information. Furthermore, Amazon has expanded into media and entertainment. It operates Prime Video, a streaming service that offers a wide range of movies, TV shows, and original content. Amazon Music provides a platform for streaming and purchasing digital music, while Audible offers audiobooks and other audio content. The company's commitment to customer satisfaction and convenience is demonstrated by its membership program, Amazon Prime. Prime members receive various benefits, including free two-day shipping, access to streaming services, exclusive deals, and more.

Industry

It services, book publishers, retail, real estate and computer and electronic product manufacturing

Company size

10,000+ Employees

Headquarters location

Seattle, WA, US