1

Overnight Kubernetes Jobs (NOW HIRING)

Sr Platform Engineer

San Jose, CA · On-site

$66.75 - $88.75/hr

Kubernetes (GKE) * Cloud Hypervisor * Linux KVM/QEMU * Open vSwitch (OVS) * OVN * Lightbits Storage * Pure Storage * Large-scale GPU infrastructure * Mixed bare-metal and virtualized environments ...

DevOps Team Lead

Indianapolis, IN · On-site

$50.50 - $69/hr

  • Medical

  • Dental

  • Vision

  • Retirement

  • PTO

... overnight travel (Travel expenses paid by Toyota Automated Logistics) Required Qualification s and ... Experience with containers and Kubernetes * Strong understanding of observability and production ...

New

DevOps Team Lead

Indianapolis, IN · On-site

$50 - $68.50/hr

  • Medical

  • Dental

  • Vision

  • Retirement

  • PTO

... overnight travel (Travel expenses paid by Toyota Automated Logistics) Required Qualification s and ... Experience with containers and Kubernetes * Strong understanding of observability and production ...

Knowledge of managed Cloud services like databases, message queues, Kubernetes clusters, docker ... Typically requires overnight travel less than 10% of the time. Education and Experience * Typically ...

Knowledge of managed Cloud services like databases, message queues, Kubernetes clusters, docker ... Typically requires overnight travel less than 10% of the time. Education and Experience * Typically ...

Storage Engineer

Manhattan, NY · On-site

$91 - $129/hr

  • Medical

  • Life

  • Retirement

Available for occasional overnight travel (10%) * Currently in possession of a valid passport or ... Familiarity with container storage solutions (Kubernetes, CSI drivers) * Vendor certifications ...

Sr. Software Engineer

Charleston, WV · Remote

$125K - $165K/yr

Travel up to 10%, including overnight travel Required Education, Experience, Certifications and ... Experience with containerization and orchestration technologies such as Docker and Kubernetes.

Sr. Software Development Engineer, MLOPs

Bellevue, WA · On-site

$138K - $182K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Key job responsibilities - Design and implement scalable ML training infrastructure on Kubernetes ... overnight, and shipping a fix to our checkpoint recovery system. After lunch, you'll pair with a ...

Cloud Software Engineer

Worthing, SD · On-site

$57.50 - $74.75/hr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

Familiarity with containerization and orchestration technologies such as Docker and Kubernetes ... Travel Travel is primarily local during the business day, although some out-of-area and overnight ...

Showing results 21-40

Overnight Kubernetes information

What cities are hiring for Overnight Kubernetes jobs?

Cities with the most Overnight Kubernetes job openings:

What are the most commonly searched types of Kubernetes jobs?

The most popular types of Kubernetes jobs are:

What states have the most Overnight Kubernetes jobs?

States with the most job openings for Overnight Kubernetes jobs include:

Sr Platform Engineer

Balin Technologies LLC

San Jose, CA • On-site

$66.75 - $88.75/hr

Other

Posted 11 days ago


Job description

Job Title: Core Platform Engineer (Infrastructure Reliability & Incident Management)
Location: Sunnyvale, CA ,  San Jose (On-Site)
Duration: Long-Term Contract
Job Description
  • As a Core Platform Engineer, we will serve as the first responder for production incidents,
  • orchestrate incident management, drive reliability improvements, and establish SRE best
  • practices across Compute, Networking, Storage, and GPU infrastructure teams.
  • Candidates should possess hands-on infrastructure experience and sufficient technical
  • depth to identify affected systems, engage the right subject matter experts, and drive
  • incident resolution processes using data and observability signals.
Responsibilities
  • Act as first responder during infrastructure incidents.
  • Lead incident bridges and coordinate cross-functional response efforts.
  • Perform incident triage and identify impacted infrastructure domains.
Gather evidence and telemetry to route incidents to the correct SME team.
  • Drive incident communications and stakeholder updates.
  • Improve reliability processes across platform engineering teams.
  • Define and promote SRE best practices and operational standards. (SLO,SLI)
  • Identify observability gaps and implement improvements.
  • Build automation for incident response workflows.
  • Manage and optimize incident management tooling (e.g., ).
  • Support change management and operational readiness processes.
  • Assist foundation engineering teams in identifying reliability risks and trends.
  • Participate in on-call activities and operational reviews.
Technical Environment
The Platform Engineering organization supports:
  • Kubernetes (GKE)
  • Cloud Hypervisor
  • Linux KVM/QEMU
  • Open vSwitch (OVS)
  • OVN
  • Lightbits Storage
  • Pure Storage
  • Large-scale GPU infrastructure
  • Mixed bare-metal and virtualized environments
Required Qualifications
  • 5 to 10+ years of experience in Site Reliability Engineering, Platform Engineering,
  • Infrastructure Operations, or Systems Engineering.
  • Strong infrastructure troubleshooting experience.
Deep expertise in at least one of the following:
  1. GPU infrastructure
  2. KVM/virtualization
  3. SDN (OVN/OVS)
  4. Storage (Lightbits/Pure Storage)
  • Proven incident management and operational leadership experience.
  • Experience running high-severity production incidents.
  • Strong understanding of observability, monitoring, SLIs, and SLOs.
  • Experience building operational automation.
  • Ability to make data-driven decisions during outages and service disruptions.
  • Preferred Qualifications
  • Experience with  or similar incident management platforms.
  • Kubernetes production operations experience.
  • Cloud-native infrastructure experience.
  • Experience supporting large-scale AI or GPU environments.
  • Strong communication and stakeholder management skills.
  • What Success Looks Like
  • Quickly identifies affected infrastructure domains during incidents.
  • Effectively coordinates SMEs and engineering teams.
  • Reduces incident response and recovery times.
  • Improves observability and operational processes.
  • Establishes reliability standards across Crusoe''s infrastructure platform.
  • Ideal Candidate Screening Criteria (For Both Roles)