Spark Recovery

60 Spark Recovery Jobs Hiring Near You

This role involves providing technical expertise for Windows Server/Desktop operating systems, administering Active Directory, and managing storage and disaster recovery solutions. The candidate will ...

Azure Databricks Architect

Seattle, WA · On-site

$72.25 - $94.25/hr

The ideal candidate will have deep expertise in Azure Databricks, Spark, Delta Lake, and modern ... Ensure scalability, reliability, security, and disaster recovery across the platform. Nice to Have

... using Spark Structured Streaming (or equivalent), including state management, checkpointing, and recovery. • Design incremental processing, partitioning strategies, and data layout/file sizing ...

Solutions Architect

Seattle, WA

$71.75 - $94.50/hr

Ensures timely recovery from outages, performs root cause analysis and implements preventative measures. * Automate, document, share, educate, and improve processes. * Develops and manages service ...

Director of Operations

Minneapolis, MN · On-site

$85K - $100K/yr

... recovery. You ensure the site gets safely back online and the team stays focused. * Build for tomorrow Partner with L&D on Spark Plug and Spark Summit readiness. Strengthen upcoming leaders through ...

Director of Operations

Blaine, MN · On-site

$85K - $100K/yr

... recovery. You ensure the site gets safely back online and the team stays focused. * Build for tomorrow Partner with L&D on Spark Plug and Spark Summit readiness. Strengthen upcoming leaders through ...

Director of Operations

Blaine, MN · On-site

$85K - $100K/yr

... recovery. You ensure the site gets safely back online and the team stays focused. * Build for tomorrow Partner with L&D on Spark Plug and Spark Summit readiness. Strengthen upcoming leaders through ...

Director of Operations

Blaine, MN · On-site

$85K - $100K/yr

... recovery. You ensure the site gets safely back online and the team stays focused. * Build for tomorrow Partner with L&D on Spark Plug and Spark Summit readiness. Strengthen upcoming leaders through ...

... recovery of domain and GPO setup • Fundamental networking concepts and troubleshooting such as DHCP, DNS, Routing, Firewall • Fundamental understanding of security principles, ACL, Encryption • ...

... Hadoop/Spark, Snowflake, etc.). - Configure and manage compute resources, including Spark ... disaster recovery, and business continuity procedures. - Automate administrative tasks and ...

... Hadoop/Spark, Snowflake, etc.). - Configure and manage compute resources, including Spark ... disaster recovery, and business continuity procedures. - Automate administrative tasks and ...

... SPARK FastTrack Award from Ann Arbor SPARK 2015 -Honoree of Diversity Focused Company by Corp! ... disaster recovery procedures Designing and/or documenting SAP procedures Provide architectural ...

Showing results 41-60

Technical Support Engineer (L3)

Galactic Minds Inc.

Chicago, IL • On-site

Other

This job post has expired today. Applications are no longer accepted.


Job description

Job Description

We are seeking an experienced Technical Support Engineer at the L3 level to provide onsite support for enterprise data engineering platforms. The successful candidate will serve as the primary technical interface at the customer location, own complex production issues through resolution, and coordinate closely with offshore engineering and support teams. This role requires deep expertise in AWS-based data services, PySpark, AWS Glue, Databricks, Amazon Redshift, and SQL, along with strong debugging, stakeholder-management, and communication skills.

Key Responsibilities

  • Provide L3 production support for data engineering applications, ETL/ELT pipelines, batch workflows, and analytical data platforms.
  • Own complex incidents from initial triage through resolution, including log analysis, data validation, defect isolation, root cause analysis, recovery, and preventive-action tracking.
  • Troubleshoot and resolve issues across AWS Glue, Amazon Redshift, Databricks, PySpark applications, SQL workloads, and related AWS services.
  • Diagnose data mismatches, pipeline failures, performance degradation, dependency issues, access problems, and environment or configuration defects.
  • Write and optimize complex SQL queries for troubleshooting, reconciliation, data-quality validation, and performance analysis.
  • Debug PySpark and Spark-based workloads using execution plans, job logs, cluster metrics, Spark UI, and application diagnostics.
  • Monitor production platforms, identify risks proactively, and ensure incidents and service requests are resolved within agreed SLAs.
  • Lead incident bridges for critical production issues and provide timely, accurate updates to technical teams, business stakeholders, and customer leadership.
  • Coordinate daily with the offshore support and data engineering teams, including work allocation, technical handoffs, follow-ups, knowledge transfer, and status reporting.
  • Partner with data engineering, platform, infrastructure, DevOps, security, QA, and vendor teams to implement permanent fixes and maintain platform stability.
  • Collaborate with different Business Units at the customer site to understand impact, prioritize issues, clarify requirements, and communicate resolution plans.
  • Support release, deployment, change, and post-production validation activities across development, test, and production environments.
  • Prepare and maintain runbooks, troubleshooting guides, known-error records, incident reports, RCA documents, and operational dashboards.
  • Identify opportunities to automate repetitive support activities, improve monitoring and alerting, and reduce manual recovery effort.
  • Mentor L1/L2 support engineers and enable the offshore team to resolve recurring issues independently.
  • Participate in an on-call or extended-hours support rotation when required for business-critical incidents.