Senior/Staff SRE for AI/ML Platform Infrastructure
San Jose, CA · On-site
$66.75 - $88.75/hr
Spot, Reserved Instances, Savings Plans. • Establishing an SRE function where one did not previously exist. Thanks & Regards, Narendra Kunware

San Jose, CA · On-site
$66.75 - $88.75/hr
Spot, Reserved Instances, Savings Plans. • Establishing an SRE function where one did not previously exist. Thanks & Regards, Narendra Kunware
San Jose, CA · On-site
$66.75 - $88.75/hr
Spot, Reserved Instances, Savings Plans. • Establishing an SRE function where one did not previously exist. Thanks & Regards, Narendra Kunware
Work closely with RFIC Design and System Engineering teams to define board requirements and provide feedback on chip-level performance. * Documentation & Data Analysis: Maintain detailed design ...
Work closely with RFIC Design and System Engineering teams to define board requirements and provide feedback on chip-level performance. * Documentation & Data Analysis: Maintain detailed design ...

$66.75 - $88.75/hr
Other
Posted 8 days ago
Minimum Qualifications
• Production on-call experience in a real rotation, with incident command and blameless postmortem practice.
• Production Kubernetes and container experience (Docker), with cloud-native infrastructure patterns.
• Hands-on production ownership on at least one major cloud (AWS, Google Cloud Platform, or Azure).
• Terraform or OpenTofu proficiency.
• Observability depth with Prometheus, Grafana, or equivalent for metrics, logging, and alerting, including dashboard and alert design.
• Strong automation skills in Python, Bash, or Go.
• Networking fundamentals: VPCs, load balancers, DNS, firewalls, cross-cloud connectivity.
• CI/CD experience with GitHub Actions, GitLab CI, Jenkins, or ArgoCD.
• Proven ability to troubleshoot complex distributed systems, largely self-directed.
Preferred Qualifications
• GPU infrastructure and AI/ML workloads: Ray, Kubeflow, MLflow, or similar.
• NVIDIA GPU orchestration: A100/H100 configuration, driver and CUDA runtime management.
• Distributed training networking: RDMA, InfiniBand, EFA, NCCL.
• Distributed tracing and OpenTelemetry instrumentation across services.
• Progressive delivery: canary and blue/green rollouts with automated rollback.
• Chaos or fault-injection testing, game days, and disaster-recovery drills.
• Multi-cloud networking, unified storage abstractions, and disaster recovery.
• FinOps and cost optimization: Spot, Reserved Instances, Savings Plans.
• Establishing an SRE function where one did not previously exist.
Thanks & Regards,
Narendra Kunware
Sourced by ZipRecruiter
It services
1,001 - 5,000 Employees
Sunnyvale, CA, US
1990