2

Remote Aiops Engineer Jobs in California (NOW HIRING)

Contributions to AIops/MLOps platforms (MLflow, Kubeflow, Vertex AI) and CI/CD for ML workflows ... Fully remote, work from home environment * Employee Share Option Plan * Flexible working hours

Remote or Fairfield, CA Job Type: Temporary (Contract) Broadcom Operations Intelligence ... Engineer automation workflows to reduce manual effort and improve signal quality. * Collaborate ...

Remote Aiops Engineer information

What is the difference between Remote Aiops Engineer vs Cloud Operations Engineer?

AspectRemote Aiops EngineerCloud Operations Engineer
Required CredentialsCertifications in AI, Machine Learning, Cloud platforms (AWS, Azure)Certifications in Cloud platforms, DevOps, Networking
Work EnvironmentRemote, tech-focused teams managing AI-driven systemsRemote or on-site, managing cloud infrastructure and services
Employer & Industry UsageTech companies, AI startups, cloud service providersCloud service providers, enterprise IT departments
Search & Comparison IntentUnderstanding AI-focused cloud operations rolesManaging cloud infrastructure and services

The Remote Aiops Engineer focuses on maintaining AI-driven systems and automating operations using AI and machine learning, often requiring specialized certifications. In contrast, the Cloud Operations Engineer manages cloud infrastructure, ensuring system reliability and performance. Both roles are remote-friendly and prevalent in tech and cloud industries, but they emphasize different technical skills and responsibilities.

What are Remote Aiops Engineers?

Remote Aiops Engineers are IT professionals who specialize in applying artificial intelligence (AI) and machine learning (ML) techniques to automate and enhance IT operations, typically while working from a location outside of a traditional office. They monitor, analyze, and optimize system performance, detect anomalies, and help prevent outages by leveraging AI-driven tools and algorithms. By working remotely, they can support organizations globally, offering flexibility and a broad skill set to manage complex IT environments efficiently.

What are the key skills and qualifications needed to thrive as a Remote AIOps Engineer, and why are they important?

To thrive as a Remote AIOps Engineer, you need a strong background in IT operations, machine learning, and data analytics, typically supported by a degree in computer science or a related field. Familiarity with monitoring tools (such as Splunk, Datadog, or Prometheus), scripting languages (like Python), and cloud platforms is essential, and certifications like AWS Certified Solutions Architect or Google Cloud Professional Engineer are highly valued. Strong problem-solving skills, effective communication, and the ability to work independently set outstanding candidates apart in a remote setting. These skills and qualifications are crucial for proactively identifying and resolving IT issues, automating processes, and ensuring system reliability in distributed environments.

How do Remote AIOps Engineers typically collaborate with cross-functional teams to resolve incidents and optimize IT operations?

Remote AIOps Engineers frequently work with IT operations, development, and security teams to monitor, analyze, and automate responses to system events and incidents. Collaboration often happens through virtual meetings, shared dashboards, and incident management tools, ensuring seamless communication despite physical distance. They play a key role in facilitating root cause analysis, sharing insights from AI-driven analytics, and implementing automation solutions to prevent future issues. Effective communication and the ability to translate technical findings for non-technical stakeholders are essential for success in this collaborative environment.
What are the most commonly searched types of Aiops Engineer jobs in California? The most popular types of Aiops Engineer jobs in California are:
What job categories do people searching Remote Aiops Engineer jobs in California look for? The top searched job categories for Remote Aiops Engineer jobs in California are:
What cities in California are hiring for Remote Aiops Engineer jobs? Cities in California with the most Remote Aiops Engineer job openings:
Infographic showing various Remote Aiops Engineer job openings in California as of July 2026, with employment types broken down into 11% Locum Tenens, 7% As Needed, 54% Full Time, 3% Part Time, 3% Contract, and 22% Nights. Highlights an 81% Physical, 6% Hybrid, and 13% Remote job distribution.

Full-time

PTO

Posted 14 days ago


Job description

Job description
Drive the end-to-end technical strategy, architecture, and productionization of ProveAI's machine learning systems, large language model (LLM) capabilities, and AI infrastructure. Own how models, evaluation pipelines, data workflows, and observability components are designed, deployed, monitored, and continuously improved to meet reliability, quality, safety, and cost goals. Provide deep AI/ML expertise and leadership across engineering teams, guiding model integration, AI/ML platform decisions, and scalable distributed systems that support enterprise-grade GenAI workloads.
Job requirements
  • 10+ years of software engineering experience with significant recent hands-on AI/ML/AI development.
  • Bachelor's degree in CS or related field.
  • Deep technical expertise in machine learning, LLMs, transformers, and modern AI frameworks (PyTorch, TensorFlow, JAX, Scikit-learn).
  • Proven experience deploying production AI/ML or LLM systems at scale (not prototypes).
  • Strong programming expertise in Python; additional experience in Java, C++, or JavaScript is a plus.
  • Experience with data engineering workflows, feature stores, and scalable data pipelines.
  • Expertise with cloud platforms (AWS/GCP/Azure), containerization, orchestration (Kubernetes), and distributed systems.
  • Hands-on AI/MLOps: model deployment, monitoring, CI/CD for AI/ML, experiment tracking, and evaluation frameworks.
  • Demonstrated technical leadership managing teams of 10+ engineers and influencing cross-functional architecture.
  • Strong ability to translate ambiguous business needs into clear technical requirements and production outcomes.
  • Expertise with LLM productionization including finetuning, retrieval-augmented generation (RAG), safety/guardrails, and evaluation.
  • Experience with AI/ML flow, Kubeflow, Vertex AI, SageMaker, or similar platforms.
  • Background in model governance, drift detection, fairness/bias evaluation, and compliance.
  • Domain specialization (NLP, computer vision, recommender systems, or agentic systems).

Nice to Have:
  • Master's or PhD in Computer Science, Machine Learning, or related discipline.
  • Cloud platform expertise (AWS, GCP, Azure) with experience deploying AI/ML workloads at scale
  • Strong product mindset with ability to translate business requirements into technical solutions
  • Contributions to AIops/MLOps platforms (MLflow, Kubeflow, Vertex AI) and CI/CD for ML workflows
  • Domain expertise in specific AI application areas such as computer vision, NLP, or recommendation systems
  • Experience with model monitoring, drift detection, and model governance in production environments
  • Previous experience with AI observability and troubleshooting

Job responsibilities
  • Define and own the architecture for scalable AI/ML systems, including training, fine-tuning, inference, evaluation, and monitoring pipelines.
  • Translate ambiguous business and product requirements into robust AI/ML system designs and staged delivery plans.
  • Make strategic decisions on model selection, LLM integrations, evaluation frameworks, model gateways, guardrails, and safety mechanisms.
  • Lead design reviews, architecture forums, and technical decision-making across teams.
  • Build and deploy production-grade AI/ML/LLM models, transformers, and generative AI features-from initial concept through production rollout.
  • Establish standards for model readiness, evaluation gates, rollout/rollback, drift detection, observability, and ongoing performance management.
  • Partner with engineering teams to integrate models into distributed systems with clear SLOs, telemetry, and error-budget mechanisms.
  • Design and improve data pipelines, feature stores, and data quality/lineage workflows supporting model training and inference.
  • Develop scalable AI/MLOps/AIOps practices for automation of training, testing, deployment, and monitoring.
  • Evaluate and implement AI/ML workflow orchestration platforms (e.g., AI/MLflow, Kubeflow, Vertex AI) and CI/CD for AI/ML.
  • Own evaluation pipelines-latency, accuracy, cost, hallucination metrics, prompt versioning, and model performance insights.
  • Instrument tracing and model observability using best-practice frameworks and telemetry standards.
  • Implement guardrails and safety systems to ensure consistent, controlled behaviour of LLM-powered features.
  • Partner closely with product, engineering, and leadership to shape platform strategy and AI feature roadmap.
  • Provide trade-off analyses that incorporate model performance, security, compliance, scalability, and long-term maintainability.
  • Write clear technical documents, proposals, and mechanism-based recommendations to guide executive decision-making.
  • Mentor senior/junior engineers in AI/ML best practices, distributed systems, experimentation, and model governance.
  • Support hiring, leveling, performance feedback, and the growth of a high-calibre engineering team.

Job benefits
  • Fully remote, work from home environment
  • Employee Share Option Plan
  • Flexible working hours
  • Paid Time-Off
  • Periodic in-person offsites globally (travel permitting)
  • Continued education support
  • Advancement opportunity