2

Remote Observability Engineer Jobs in New Mexico

Senior AI Systems Engineer

Albuquerque, NM · On-site +1

$95K - $130K/yr

Maintain observability across AI systems through logging, metrics, performance monitoring, alerting ... This position may be performed fully remote, hybrid, or onsite at an ARA office. Preference will be ...

Remote Observability Engineer information

What is a remote observability engineer?

A Remote Observability Engineer is a professional responsible for designing, implementing, and maintaining systems that monitor the health, performance, and reliability of software applications and infrastructure from a remote location. They use observability tools to collect and analyze logs, metrics, and traces, helping organizations quickly detect and resolve issues. Their work ensures that distributed systems are transparent, reliable, and efficient, often collaborating with development, operations, and security teams. Remote Observability Engineers often work from anywhere, leveraging cloud-based tools and platforms to manage complex IT environments.

What are the typical collaboration patterns for a remote observability engineer working with distributed teams?

Remote Observability Engineers frequently collaborate with software developers, DevOps teams, and IT operations to ensure systems are monitored effectively and issues are detected early. Working remotely, you'll often use communication tools like Slack, Jira, and video conferencing to coordinate incident response, discuss monitoring strategies, and review system health dashboards. Regular sync meetings and asynchronous updates are common, and you'll likely contribute to documentation and knowledge sharing to keep all stakeholders informed. Building strong communication habits is important, as much of the troubleshooting and improvement work hinges on clear coordination with multiple teams.

What are the key skills and qualifications needed to thrive as a remote observability engineer, and why are they important?

To thrive as a Remote Observability Engineer, you need strong expertise in monitoring, logging, and tracing systems, along with a background in computer science or related technical fields. Familiarity with tools like Prometheus, Grafana, ELK Stack, Datadog, and cloud platforms is typically required, as well as relevant certifications such as AWS Certified Cloud Practitioner or Google Cloud Professional DevOps Engineer. Excellent problem-solving abilities, communication skills, and a proactive mindset help you detect and resolve issues before they impact users. These competencies ensure system reliability, enable rapid incident response, and support seamless collaboration in distributed environments.

What is the difference between Remote Observability Engineer vs Site Reliability Engineer?

AspectRemote Observability EngineerSite Reliability Engineer
CredentialsKnowledge of monitoring tools, scripting, cloud platformsSame as Observability Engineer, plus SRE certifications often preferred
Work EnvironmentFocus on monitoring, logging, and tracing systems remotelyBroader scope including system reliability, incident response, and automation
Industry UsagePrimarily in tech, SaaS, cloud servicesWidely in tech, finance, and large-scale online services

The Remote Observability Engineer specializes in monitoring and analyzing system performance remotely, focusing on tools like logs and metrics. In contrast, the Site Reliability Engineer has a broader role, ensuring overall system reliability, automation, and incident management. While both roles require similar technical skills, SREs often have additional responsibilities related to system resilience and scalability.

What are the most commonly searched types of Observability Engineer jobs in New Mexico?

The most popular types of Observability Engineer jobs in New Mexico are:

What are popular job titles related to Remote Observability Engineer jobs in New Mexico?

For Remote Observability Engineer jobs in New Mexico, the most frequently searched job titles are:

What job categories do people searching Remote Observability Engineer jobs in New Mexico look for?

The top searched job categories for Remote Observability Engineer jobs in New Mexico are:

Senior Site Reliability Engineer

Berriehill Research

Albuquerque, NM • On-site, Remote

$52 - $69.25/hr

Full-time

Re-posted 2 days ago


Job description

Essential Functions:

  • Partner with software developers, platform engineers, and IT staff to improve system design, operability, deployment safety, and production support readiness.
  • Define and maintain operational standards, runbooks, support procedures, escalation paths, and service-level objectives.
  • Evaluate system architecture and changes to ensure they balance functional requirements, service quality, reliability, security, and compliance needs.
  • Drive continuous improvement in platform stability, maintenance, and availability.
  • Provide advanced technical support and troubleshooting for complex platform and service issues affecting internal users and stakeholders.

Experience and Skills Required:

  • 8+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Systems Engineering, or related infrastructure roles supporting production services.
  • Strong experience with Linux systems administration and troubleshooting in enterprise environments.
  • Strong experience operating and maintaining on-prem Kubernetes platforms and all related components including CRI, CNI, and CSI plugins.
  • Experience deploying and maintaining applications on Kubernetes using Helm, Kustomize, and similar tooling.
  • Experience supporting DevOps tooling such as GitLab, Artifactory, Jira, Confluence.
  • Experience with GitOps tools such as FluxCD or ArgoCD.
  • Proficiency scripting with at least one of Python, Go, or Bash.
  • Strong experience designing, maintaining, and maturing observability tooling including monitoring, dashboards, logging and tracing, and supporting SLOs.
  • Strong understanding of reliability engineering concepts:
    • Service health indicators
    • High availability design, failure reduction, and testing
    • Operational readiness practices, including developing documentation, runbooks, and architectural descriptions
    • Incident response, root cause analysis, remediation/recovery
  • Ability to obtain a security clearance, which includes U.S. citizenship.

Preferred:

  • Experience with multiple Linux distributions including Ubuntu.
  • Experience with at least one of the following: Tanzu Kubernetes, Nutanix Kubernetes Platform, Canonical Kubernetes.
  • Experience with cloud platforms such as AWS and Azure.
  • Experience with infrastructure automation and configuration management.
  • Experience managing AI tooling on Kubernetes including MCP Servers, LLM platforms (vLLM, Ollama), Kubeflow.
  • Experience with security and compliance considerations in regulated environments.
  • DoD experience.
  • Active or inactive Secret Security Clearance.

Education:

  • Bachelor’s degree in CS, Software Engineering or other IT-related field or equivalent experience

REMOTE WORK NOTICE:  This position may be performed fully remote, hybrid, or onsite at an ARA office. Preference will be given to candidates located onsite in the Albuquerque area.