1

Home Based Observability Engineer Jobs (NOW HIRING)

... based telemetry ingestion for logs, metrics, traces, and spans across distributed systems. • ... , platform and business teams to embed proactive intelligence and Observability standards ...

New

As a leading financial services and healthcare technology company based on revenue, SS&C is ... Job Title: Sr. Observability Engineer Locations: [Waltham, MA - Hybrid] About the Role We are ...

As a leading financial services and healthcare technology company based on revenue, SS&C is ... Observability Engineer (Splunk) Locations: Jacksonville, FL; Boston, MA; Kanas City, MO | Hybrid 6x ...

The Director will start with two direct reports - a to be hired Senior Observability Architect (India-based) and a future US-based Observability/Reliability Engineer - and will be expected to scale ...

As a leading financial services and healthcare technology company based on revenue, SS&C is ... Observability Engineer (Splunk) Locations: Jacksonville, FL; Boston, MA; Kanas City, MO | Hybrid 6x ...

As a leading financial services and healthcare technology company based on revenue, SS&C is ... Observability Engineer (Splunk) Locations: Jacksonville, FL; Boston, MA; Kanas City, MO | Hybrid 6x ...

Lead Observability Engineer (Grafana Cloud)

Coppell, TX · On-site

$95K - $125K/yr

Comprehensive health and life insurance and well-being benefits, based on location * Pension ... Being a member of IT FinSight Delivery team, you will be for an Observability Engineer that will be ...

As a leading financial services and healthcare technology company based on revenue, SS&C is ... Observability Engineer (Splunk) Locations: Jacksonville, FL; Boston, MA; Kanas City, MO | Hybrid 6x ...

Senior AI and HPC Observability Engineer

Seattle, WA · On-site

$139K - $183K/yr

We are looking for a strong AI & HPC Observability Engineer to build and scale next-generation ... will be determined based on your location, experience, and the pay of employees in similar ...

Lead Observability Engineer (Grafana Cloud)

Tampa, FL · On-site

$93K - $122K/yr

Comprehensive health and life insurance and well-being benefits, based on location * Pension ... Being a member of IT FinSight Delivery team, you will be for an Observability Engineer that will be ...

next page

Showing results 1-20

Home Based Observability Engineer information

What is the difference between Home Based Observability Engineer vs Network Operations Center (NOC) Technician?

AspectHome Based Observability EngineerNetwork Operations Center (NOC) Technician
CredentialsRelevant certifications like Cisco CCNA, CompTIA Network+Similar certifications often required, such as Cisco CCNA
Work EnvironmentRemote, home-based setup with monitoring toolsOn-site or remote, monitoring network systems in a control room
Industry UsageIT, cloud services, software companiesTelecommunications, internet service providers
Job FocusMonitoring, troubleshooting, and optimizing observability toolsNetwork monitoring, incident response, and system maintenance

While both roles involve monitoring network and system health, the Home Based Observability Engineer focuses on software observability tools and cloud environments remotely, whereas the NOC Technician primarily manages network infrastructure on-site or remotely. Both require similar certifications and are vital in maintaining system uptime, but their daily tasks and work settings differ.

More about Home Based Observability Engineer jobs
What cities are hiring for Home Based Observability Engineer jobs? Cities with the most Home Based Observability Engineer job openings:
What are the most commonly searched types of Observability Engineer jobs? The most popular types of Observability Engineer jobs are:
What states have the most Home Based Observability Engineer jobs? States with the most job openings for Home Based Observability Engineer jobs include:
What job categories do people searching Home Based Observability Engineer jobs look for? The top searched job categories for Home Based Observability Engineer jobs are:
Infographic showing various Home Based Observability Engineer job openings in the United States as of July 2026, with employment types broken down into 1% As Needed, 78% Full Time, 14% Part Time, and 7% Contract. Highlights an 92% Physical, 2% Hybrid, and 6% Remote job distribution.
AI/ML Observability Engineer

AI/ML Observability Engineer

StradIT

Dallas, TX • On-site

Contractor

Posted 2 days ago


Job description

VISA INDEPENDENT CANDIDATES ONLY !
Overview
We are seeking a passionate and hands-on AI/ML Engineer to accelerate our Enterprise Observability strategy. This role will design, build, and operationalize AI/ML capabilities that enhance end to end telemetry pipelines, anomaly detection, intelligent alerting, and proactive system resiliency.
You will work at the intersection of AI/ML engineering, Observability platforms, and automation, developing solutions that improve detection, diagnosis, and prevention of operational issues across distributed systems.
Requirements
Key Responsibilities
• Design and deploy AI/ML models supporting anomaly detection, baselining, event correlation, and predictive operational analytics.
• Build and integrate AI-enabled capabilities into enterprise Observability platforms, including Grafana, APM/RUM tools, network telemetry systems, and data observability tools.
• Develop AI Agents that can autonomously triage issues, recommend corrective actions, and initiate automated remediation workflows to reduce recovery time and improve system resilience.
• Implement self-healing automation using AI-driven decisioning, integrating with orchestration frameworks, service APIs, and infrastructure automation pipelines.
• Engineer and maintain real-time and batch data pipelines using Snowflake ML Jobs, Snowflake Cortex, streams, tasks, and UDFs.
• Implement and manage OpenTelemetry-based telemetry ingestion for logs, metrics, traces, and spans across distributed systems.
• Build asynchronous Python APIs and services for model inferencing and operational integration.
• Enhance observability intelligence with AI-powered capabilities such as root-cause acceleration, chatbot/search enablement, and automated insights.
• Contribute to SLO/SLI modeling, Golden Signals instrumentation, and Observability NFR adoption.
• Collaborate across engineering, SRE, platform and business teams to embed proactive intelligence and Observability standards throughout the ecosystem.
Required Skills & Qualifications
Core Technical Skills
• Strong proficiency in Python and data science/ML libraries:
NumPy, Pandas, scikit learn, TensorFlow, PyTorch, Matplotlib, Seaborn.
• Experience with Generative AI, LLM fine tuning, prompt engineering, RAG pipelines, and LLM evaluation frameworks.
• Expertise in developing and deploying ML models in production (batch & streaming).
• Strong understanding of statistics, time series modeling, and anomaly detection.
Observability & Telemetry
• Experience with OpenTelemetry for logs, metrics, traces, spans.
• Familiarity with Observability concepts:
Golden Signals, SLO/SLI design, APM, RUM, Synthetics, event correlation, baselining.
• Experience with Observability tools such as:
Grafana (Alloy agents, dashboards, ML capabilities), Dynatrace, Monte Carlo (Data Observability), Netscout, ThousandEyes, SolarWinds, NetBrain.
Cloud, Data & Platform
• Hands on with AWS (SageMaker, Bedrock), Snowflake ML, Snowflake/Openflow, Snowflake AI Observability tooling.
• Experience building Snowflake data pipelines (streams, tasks, UDFs) - plus for Cortex features.
• Strong understanding of distributed systems and microservices telemetry requirements.
Automation & Engineering Quality
• Experience with automation pipelines, CI/CD, and infrastructure as code patterns supporting Observability adoption.
• Ability to build asynchronous Python APIs or services for model inference and operational integration.
Preferred Qualifications
• Experience developing agentic AI systems that analyze telemetry, generate action recommendations, or execute automated operational responses.
• Experience building self-healing patterns, including automated rollback, service restarts, configuration corrections, and predictive maintenance.
• Experience in Snowflake ML workflows, Snowflake Cortex Agents, and data pipeline automation.
• Exposure to AI-enabled alerting, RCA automation, and operational self-healing concepts.
• Experience with large-scale operational telemetry and multi-cloud ecosystems.
Soft Skills
• Strong analytical thinking and problem solving.
• Excellent communication skills for cross functional collaboration with infrastructure, SRE, engineering, business, and leadership teams.
• Curiosity, continuous learning mindset, and passion for applied AI and Observability.