2

Remote Kubeflow Jobs in California (NOW HIRING)

Experience with AI/ML flow, Kubeflow, Vertex AI, SageMaker, or similar platforms. * Background in ... Fully remote, work from home environment * Employee Share Option Plan * Flexible working hours

Sr. ML Ops Engineer

Mountain View, CA · On-site +1

$123K - $169K/yr

... hybrid or remote role with periodic trips to HQ in Mountain View, CA. Must Haves * 2-3 years ... Background in robotics autonomy and computer vision Experience integrating with tools like Kubeflow ...

Senior Engineer - LLMOps & MLOps

Los Angeles, CA · On-site +1

$112K - $154K/yr

... Kubeflow, or Step Functions). LLM Tooling: Professional experience with evaluation and ... remote #LI-TS1 Sedgwickis an Equal Opportunity Employer and a Drug-Free Workplace. If you're ...

Remote Kubeflow information

What is the difference between Remote Kubeflow vs Remote Data Scientist?

AspectRemote KubeflowRemote Data Scientist
Required CredentialsCloud certifications, Kubernetes, ML OpsStatistics, Machine Learning, Programming
Work EnvironmentCloud platforms, DevOps toolsData analysis, modeling, research
Industry UsageAI/ML deployment, MLOps teamsData analysis, predictive modeling

Remote Kubeflow focuses on deploying and managing ML workflows using Kubernetes, requiring cloud and DevOps skills. Remote Data Scientists analyze data, build models, and interpret results. While both roles involve machine learning, Remote Kubeflow emphasizes deployment and infrastructure, whereas Remote Data Scientists focus on data analysis and modeling.

What are the key skills and qualifications needed to thrive as a remote Kubeflow engineer?

To thrive as a Remote Kubeflow Engineer, you need strong expertise in machine learning, cloud computing, and container orchestration, typically supported by a degree in computer science or related fields. Proficiency with tools such as Kubeflow, Kubernetes, Docker, and cloud platforms like AWS, GCP, or Azure—as well as experience with CI/CD pipelines—is essential. Strong problem-solving skills, communication, and the ability to collaborate remotely are important soft skills for success. These skills ensure the effective deployment and management of scalable machine learning workflows in distributed, cloud-based environments.

What are some common challenges faced by professionals working in a remote Kubeflow engineer role?

Remote Kubeflow engineers often encounter challenges such as troubleshooting distributed machine learning pipelines without direct, on-premises access to infrastructure. Effective communication with data scientists, DevOps, and other stakeholders can also be more complex due to differing time zones and remote collaboration tools. Additionally, managing secure access and ensuring seamless deployment of ML workflows in cloud environments requires a strong understanding of both Kubernetes and Kubeflow. Overcoming these challenges typically involves proactive documentation, regular virtual meetings, and a collaborative approach to problem-solving.

What is a remote Kubeflow?

A Remote Kubeflow job refers to a role where professionals use Kubeflow, an open-source machine learning platform designed for Kubernetes, while working remotely. These jobs typically involve designing, deploying, and managing machine learning workflows on cloud or on-premises Kubernetes clusters. Responsibilities may include automating ML pipelines, optimizing model training, and collaborating with data scientists and engineers. Remote Kubeflow professionals usually need expertise in Kubernetes, Docker, Python, and machine learning concepts. The remote aspect allows them to perform these tasks from anywhere with reliable internet access.

What are the most commonly searched types of Kubeflow jobs in California?

The most popular types of Kubeflow jobs in California are:

What job categories do people searching Remote Kubeflow jobs in California look for?

The top searched job categories for Remote Kubeflow jobs in California are:

What cities in California are hiring for Remote Kubeflow jobs?

Cities in California with the most Remote Kubeflow job openings:

Full-time

PTO

Re-posted 3 days ago


Job description

Job description
Drive the end-to-end technical strategy, architecture, and productionization of ProveAI's machine learning systems, large language model (LLM) capabilities, and AI infrastructure. Own how models, evaluation pipelines, data workflows, and observability components are designed, deployed, monitored, and continuously improved to meet reliability, quality, safety, and cost goals. Provide deep AI/ML expertise and leadership across engineering teams, guiding model integration, AI/ML platform decisions, and scalable distributed systems that support enterprise-grade GenAI workloads.
Job requirements
  • 10+ years of software engineering experience with significant recent hands-on AI/ML/AI development.
  • Bachelor's degree in CS or related field.
  • Deep technical expertise in machine learning, LLMs, transformers, and modern AI frameworks (PyTorch, TensorFlow, JAX, Scikit-learn).
  • Proven experience deploying production AI/ML or LLM systems at scale (not prototypes).
  • Strong programming expertise in Python; additional experience in Java, C++, or JavaScript is a plus.
  • Experience with data engineering workflows, feature stores, and scalable data pipelines.
  • Expertise with cloud platforms (AWS/GCP/Azure), containerization, orchestration (Kubernetes), and distributed systems.
  • Hands-on AI/MLOps: model deployment, monitoring, CI/CD for AI/ML, experiment tracking, and evaluation frameworks.
  • Demonstrated technical leadership managing teams of 10+ engineers and influencing cross-functional architecture.
  • Strong ability to translate ambiguous business needs into clear technical requirements and production outcomes.
  • Expertise with LLM productionization including finetuning, retrieval-augmented generation (RAG), safety/guardrails, and evaluation.
  • Experience with AI/ML flow, Kubeflow, Vertex AI, SageMaker, or similar platforms.
  • Background in model governance, drift detection, fairness/bias evaluation, and compliance.
  • Domain specialization (NLP, computer vision, recommender systems, or agentic systems).

Nice to Have:
  • Master's or PhD in Computer Science, Machine Learning, or related discipline.
  • Cloud platform expertise (AWS, GCP, Azure) with experience deploying AI/ML workloads at scale
  • Strong product mindset with ability to translate business requirements into technical solutions
  • Contributions to AIops/MLOps platforms (MLflow, Kubeflow, Vertex AI) and CI/CD for ML workflows
  • Domain expertise in specific AI application areas such as computer vision, NLP, or recommendation systems
  • Experience with model monitoring, drift detection, and model governance in production environments
  • Previous experience with AI observability and troubleshooting

Job responsibilities
  • Define and own the architecture for scalable AI/ML systems, including training, fine-tuning, inference, evaluation, and monitoring pipelines.
  • Translate ambiguous business and product requirements into robust AI/ML system designs and staged delivery plans.
  • Make strategic decisions on model selection, LLM integrations, evaluation frameworks, model gateways, guardrails, and safety mechanisms.
  • Lead design reviews, architecture forums, and technical decision-making across teams.
  • Build and deploy production-grade AI/ML/LLM models, transformers, and generative AI features-from initial concept through production rollout.
  • Establish standards for model readiness, evaluation gates, rollout/rollback, drift detection, observability, and ongoing performance management.
  • Partner with engineering teams to integrate models into distributed systems with clear SLOs, telemetry, and error-budget mechanisms.
  • Design and improve data pipelines, feature stores, and data quality/lineage workflows supporting model training and inference.
  • Develop scalable AI/MLOps/AIOps practices for automation of training, testing, deployment, and monitoring.
  • Evaluate and implement AI/ML workflow orchestration platforms (e.g., AI/MLflow, Kubeflow, Vertex AI) and CI/CD for AI/ML.
  • Own evaluation pipelines-latency, accuracy, cost, hallucination metrics, prompt versioning, and model performance insights.
  • Instrument tracing and model observability using best-practice frameworks and telemetry standards.
  • Implement guardrails and safety systems to ensure consistent, controlled behaviour of LLM-powered features.
  • Partner closely with product, engineering, and leadership to shape platform strategy and AI feature roadmap.
  • Provide trade-off analyses that incorporate model performance, security, compliance, scalability, and long-term maintainability.
  • Write clear technical documents, proposals, and mechanism-based recommendations to guide executive decision-making.
  • Mentor senior/junior engineers in AI/ML best practices, distributed systems, experimentation, and model governance.
  • Support hiring, leveling, performance feedback, and the growth of a high-calibre engineering team.

Job benefits
  • Fully remote, work from home environment
  • Employee Share Option Plan
  • Flexible working hours
  • Paid Time-Off
  • Periodic in-person offsites globally (travel permitting)
  • Continued education support
  • Advancement opportunity