1

International Kubernetes Jobs (NOW HIRING)

Senior DevOps Engineer

$133K - $170K/yr

... international companies building scalable cloud-native platforms and enterprise infrastructure. This role is ideal for engineers with deep expertise in Kubernetes, CI/CD automation, cloud ...

DevOps Engineer

San Jose, CA · On-site

$60.50 - $82.75/hr

Fast moving, challenging and unique business problems International work environment and flat ... Kubernetes arch and component, have good experience in large-scale cluster operation and ...

Senior Data Engineer

Olathe, KS · On-site

$100K - $136K/yr

Administers/maintains a container-based and Kubernetes-based Airflow installations * Understands ... Garmin International is an equal opportunity employer. Qualified applicants will receive ...

Description & Requirements Elevate your career with MANTECH International Corporation! Join a ... Maintain, optimize, and troubleshoot production environments utilizing Kubernetes, OpenShift, and ...

Senior DevOps Engineer

Las Vegas, NV · On-site

$124K - $159K/yr

Arrow International is the world's #1 maker of charitable gaming products, from pull tabs and bingo ... Manage and optimize our containerized environments using Docker and Kubernetes. * Implement and ...

Senior DevOps Engineer

Las Vegas, NV · On-site

$124K - $159K/yr

Join the Fun at Arrow International! Arrow International is the world's #1 maker of charitable ... Manage and optimize our containerized environments using Docker and Kubernetes. * Implement and ...

Site Reliability Engineer

Plano, TX · On-site

$54.50 - $72.50/hr

Openshift Kubernetes Development Experience(Java, Python, Golang) SRE Skills Nice to Haves: Baremetal Cloud Client We are looking for a highly skilled Site Reliability and operations Engineer (SRE) ...

Showing results 21-40

International Kubernetes information

Senior Site Reliability Engineer

Castleton Commodities International

Stamford, CT • On-site

$60.75 - $80.75/hr

Full-time

This job post has expired 1 day ago. Applications are no longer accepted.


Job description

Job Summary:
Castleton Commodities International is seeking a Senior Site Reliability Engineer responsible for improving the reliability, availability, scalability, and operational excellence of critical infrastructure platforms and services. The role involves partnering with Engineering, Security, and Infrastructure teams to design resilient architectures and implement Infrastructure as Code and CI/CD standards.
Responsibilities:
• Own and improve service reliability through SLO/SLI definition, error budgets, and operational best practices.
• Design, implement, and maintain observability (monitoring, logging, tracing, alerting) to reduce MTTR and improve proactive detection.
• Lead incident response practices including on-call improvements, runbooks, post-incident reviews (RCA), and preventative actions.
• Partner with application teams to improve performance, capacity planning, and resiliency under failure scenarios.
• Design and operate highly available, fault-tolerant Cloud architectures (multi-AZ and, where required, multi-region).
• Implement resilient patterns across compute, storage, networking, and managed services (e.g., autoscaling, load balancing, backups, replication).
• Drive cloud governance best practices (tagging, account/landing zone patterns, least privilege, guardrails) in partnership with security and platform teams.
• Build and maintain IaC modules and standards (e.g., Terraform, CloudFormation, CDK) for repeatable, auditable infrastructure delivery.
• Develop, standardize, and optimize CI/CD pipelines to enable safe, automated deployments (e.g., GitHub Actions, GitLab CI, Jenkins, AWS CodePipeline).
• Promote DevOps practices: version-controlled infrastructure, automated testing, immutable deployments, and progressive delivery patterns.
• Establish environment consistency across dev/test/stage/prod and ensure infrastructure drift detection and remediation.
• Collaborate with stakeholders to evaluate and define service-level RTO and RPO targets based on business and technical requirements.
• Design and implement BCP/DR architectures and procedures (backups, restore workflows, replication, failover/failback, data integrity validation).
• Coordinate and execute structured DR tests (tabletop, simulation, partial failover, full failover) and document outcomes.
• Maintain DR runbooks, dependency maps, and recovery checklists; drive remediation of gaps identified during testing.
• Produce metrics and reporting on DR readiness, test results, and continuous improvement actions.
Qualifications:
Required:
• 7+ years of experience in SRE, DevOps, Platform Engineering, or Systems Engineering roles supporting production environments.
• Strong proficiency with observability platforms (e.g., Datadog, Prometheus/Grafana, ELK/OpenSearch, Nagios, Nimsoft, etc).
• Strong hands-on AWS experience building and operating production systems.
• Proven expertise with Infrastructure as Code (Terraform and/or CloudFormation/CDK).
• Strong CI/CD and automation background (pipeline design, deployment strategies, testing automation).
• Experience defining and validating RTO/RPO, and implementing BCP/DR plans with structured testing.
• Experience with Kubernetes and auto-scaling container platforms (EKS, ECS, or Kubernetes on-prem).
• Strong Linux fundamentals, networking concepts (DNS, TCP/IP, load balancing), and troubleshooting skills.
• Proficiency in at least one scripting/programming language (Python, Go, Bash, or similar).
• Ability to write clear operational documentation, runbooks, and post-incident reports.
• Ability to work effectively in a fast-paced, dynamic and high-intensity environment including open-floor plan if applicable to the position, with timely responsiveness and the ability to work beyond normal business hours when required.
Preferred:
• Familiarity with Azure and/or Oracle Cloud (OCI).
• Familiarity with Service Mesh, API Gateways, and distributed tracing tooling.
• Familiarity with OpenTelemetry, client instrumentations and collector configurations.
• Security and compliance familiarity in cloud environments (IAM design, secrets management, audit logging).
• Experience implementing progressive delivery (blue/green, canary), feature flags, and automated rollback.
• Relevant certifications (AWS Solutions Architect/DevOps Engineer, Kubernetes CKA/CKAD).
• Experience with ArgoCD & Karpenter.
Company:
Castleton Commodities International is an independent global commodities merchant. Founded in 2001, the company is headquartered in Stamford, USA, with a team of 201-500 employees. The company is currently Growth Stage.