2

Remote Observability Jobs in Utah (NOW HIRING)

$56.50 - $75.25/hr

... remote state on S3. * strong> CI/CD & automation: pipeline design in GitHub Actions or GitLab CI ... Observability: Datadog, OpenTelemetry (collector and kube-stack), Grafana; alerting hygiene with ...

This position is remote eligible for candidates who currently reside in Utah. What You'll Do ... Observability: Leverage Datadog and other monitoring tools to implement "Shift-Right" testing ...

Software Engineer - AI-Native Full Stack Bolo AI Bay Area (Hybrid) | Salt Lake City Area (Remote ... Reliability, observability, and graceful degradation matter here. What Makes Someone Good at This ...

This position is remote eligible for candidates who currently reside in Utah. Click here to see why ... Experience with observability tooling (logging, metrics, tracing) * Hands-on experience building ...

This position is remote eligible for candidates who currently reside in Utah. What You'll Do ... Observability: Leverage Datadog and other monitoring tools to implement "Shift-Right" testing ...

Showing results 21-26

Remote Observability information

What are some common challenges faced by professionals in a remote observability role, and how can they be addressed?

Professionals in Remote Observability often face challenges such as monitoring complex, distributed systems, ensuring reliable data collection, and quickly identifying the root causes of issues without physical access to infrastructure. To address these challenges, it's essential to implement robust monitoring tools, establish clear alerting thresholds, and maintain strong communication with development and operations teams. Regular knowledge-sharing sessions and continuous learning about new observability platforms can also help remote teams stay effective and proactive.

What is the difference between Remote Observability vs Remote Monitoring?

AspectRemote ObservabilityRemote Monitoring
FocusComprehensive system insights, including logs, metrics, and tracesTracking specific system metrics and alerts
ToolsOpenTelemetry, Grafana, JaegerNagios, Zabbix, Datadog
Work EnvironmentDevOps, SRE teams managing complex distributed systemsIT operations teams overseeing system health
CredentialsKnowledge of cloud platforms, scripting, and monitoring toolsBasic networking, system administration skills

Remote Observability provides a holistic view of system health through logs, metrics, and traces, enabling proactive troubleshooting. Remote Monitoring focuses on tracking specific metrics and alerts to detect issues. While both roles involve system oversight, observability offers deeper insights for complex environments, whereas monitoring emphasizes real-time alerts for system stability.

What are the key skills and qualifications needed to thrive as a remote observability engineer?

To thrive as a Remote Observability Engineer, you need expertise in monitoring, logging, and tracing, typically supported by experience in systems administration or DevOps and a relevant technical degree. Familiarity with observability tools like Prometheus, Grafana, Datadog, ELK Stack, and cloud monitoring platforms, as well as certifications such as AWS Certified Cloud Practitioner or Google Professional Cloud DevOps Engineer, is highly valued. Strong analytical thinking, problem-solving, and effective communication are vital soft skills for diagnosing issues and collaborating with distributed teams. These skills and qualifications ensure reliable system performance, rapid incident response, and seamless user experiences in complex, cloud-based environments.

What is remote observability?

Remote observability refers to the ability to monitor, measure, and understand the state and performance of systems, applications, or infrastructure from a distance, typically using specialized tools and platforms. It is crucial for organizations that operate distributed or cloud-based environments, as it allows teams to detect issues, analyze metrics, and ensure reliability without needing physical access to the hardware. Remote observability often involves collecting logs, metrics, traces, and other telemetry data to provide a comprehensive view of system health and performance.
What are the most commonly searched types of Observability jobs in Utah? The most popular types of Observability jobs in Utah are:
What job categories do people searching Remote Observability jobs in Utah look for? The top searched job categories for Remote Observability jobs in Utah are:
What cities in Utah are hiring for Remote Observability jobs? Cities in Utah with the most Remote Observability job openings:

$56.50 - $75.25/hr

Full-time

Medical

Posted 8 days ago


Job description

Cloud Platform Engineer (AWS / EKS / Terraform)
Team: Infrastructure · Reports to Head of Infrastructure
About the role
You will build and operate TookiTaki's cloud platform: AWS EKS clusters managed as code, GitOps-driven delivery with the Argo suite, operator-managed data services running on Kubernetes, and a Terraform-based self-service layer that lets application teams request infrastructure through reviewed YAML instead of tickets. The job is platform engineering, not click-ops — everything ships through version control, automated pipelines, and policy gates.
Requirements
Education

  • R
  • equired: Bachelor's degree in Computer Science, Engineering, or a related field — or equivalent practical experience.
  • P
  • referred: Relevant certifications over a Master's — CKA (Certified Kubernetes Administrator), HashiCorp Terraform Associate, AWS Solutions Architect Associate or higher.
    Experience
  • 3
  • –5+ years in cloud, platform, or DevOps engineering.
  • P
  • roven track record running production Kubernetes on a managed cloud (EKS strongly preferred), including stateful workloads.
  • E
  • xperience operating infrastructure entirely through Infrastructure as Code — no console-driven change management.
    Technical expertise
  • <
  • strong>Kubernetes / EKS: cluster lifecycle and upgrades, autoscaling (Karpenter, KEDA, VPA), ingress and load balancing, IRSA/Pod Identity, core add-ons (cert-manager, external-dns, CoreDNS, node-local-dns).
  • <
  • strong>GitOps / Argo: ArgoCD for platform and application delivery (app-of-apps, Helm chart authoring, sync and rollback strategies), Argo Rollouts for progressive delivery, Argo Workflows/Events for automation.
  • <
  • strong>Databases and stateful services on Kubernetes: deploying and operating databases via Kubernetes operators — PostgreSQL and MySQL (Percona operators), ScyllaDB, Elasticsearch, Valkey/Redis, Kafka (Strimzi) — plus AWS RDS/Aurora where managed services fit better; backup/restore, upgrades, and capacity management for stateful workloads.
  • <
  • strong>Terraform: authoring reusable, versioned, tested modules (not just consuming them) — variable/output interface design, `terraform test`, semantic versioning, remote state on S3.
  • <
  • strong>CI/CD & automation: pipeline design in GitHub Actions or GitLab CI; plan/apply automation with Atlantis; policy-as-code gates (OPA/Conftest, checkov, tflint); pre-merge validation and drift detection.
  • <
  • strong>AWS core services: VPC and network design, Route 53, IAM (least-privilege roles and policies), S3, RDS/Aurora.
  • <
  • strong>Observability: Datadog, OpenTelemetry (collector and kube-stack), Grafana; alerting hygiene with low false-positive rates; log and event pipelines.
  • <
  • strong>Programming: solid scripting/tooling ability in Python or Go — enough to build renderers, validators, and pipeline tooling, not just glue scripts.
  • <
  • strong>Nice to have: data platform tooling (Airflow, Spark on Kubernetes, Temporal, StarRocks), identity and SSO (Keycloak, Dex, oauth2-proxy), load testing with k6, FinOps practices (tagging standards, cost allocation, rightsizing).
    Soft skills
  • S
  • trong problem-solving and analytical abilities; comfortable debugging across the stack (DNS → LB → cluster → workload → database).
  • C
  • lear written communication — design docs, runbooks, and PR descriptions are first-class deliverables here.
  • C
  • ollaborative mindset: infrastructure changes ship through peer review, and platform decisions are made with (not for) application teams.
    Key competencies
  • <
  • strong>Platform thinking: design paved roads that make the secure, cost-efficient path the easy path for application teams.
  • <
  • strong>Automation-first: if a task is done twice manually, the third time is a pipeline.
  • <
  • strong>Ownership: own services end to end — provisioning, upgrades, incidents, cost, and documentation.
  • <
  • strong>Cost awareness: treat cloud spend as an engineering metric; tag, measure, and optimize continuously.
  • <
  • strong>Adaptability: comfortable in a fast-moving environment where the platform itself is under active development.
    Success metrics (first 6–12 months)
  • A
  • pplication teams provision standard infrastructure through the self-service platform with no manual Terraform written by requesters.
  • E
  • KS cluster and node-group upgrades executed as routine, zero-downtime operations via blue/green rollout.
  • O
  • perator-managed data services (PostgreSQL, Kafka, Elasticsearch, etc.) run with tested backup/restore and rehearsed upgrade procedures.
  • 1
  • 00% of infrastructure changes delivered through reviewed, policy-gated pipelines — zero out-of-band console changes.
  • D
  • eployment lead time for platform changes reduced measurably (target: same-day merge-to-production for routine changes).
  • C
  • ost visibility established via tagging and dashboards, with identified savings executed (rightsizing, autoscaling, storage tiering).
  • A
  • ctionable alerting: on-call pages correspond to real incidents; false-positive alerts trend toward zero.
  • M
  • ean time to recovery for platform incidents under 30 minutes.
    Benefits
  • <
  • strong>Competitive salary aligned with industry standards and experience.
  • <
  • strong>Professional development: certification support (CKA, AWS, Terraform) and training across cloud, platform, and data engineering.
  • <
  • strong>Comprehensive benefits: health insurance and flexible working options.
  • <
  • strong>Growth opportunities: career progression within TookiTaki's expanding infrastructure and platform organization.

    Employment Type: FULL_TIME