Debug PySpark and Spark-based workloads using execution plans, job logs, cluster metrics, Spark UI, ... manual recovery effort. * Mentor L1/L2 support engineers and enable the offshore team to resolve ...
60 Spark Recovery Jobs Hiring Near You
Debug PySpark and Spark-based workloads using execution plans, job logs, cluster metrics, Spark UI, ... manual recovery effort. * Mentor L1/L2 support engineers and enable the offshore team to resolve ...
This role involves providing technical expertise for Windows Server/Desktop operating systems, administering Active Directory, and managing storage and disaster recovery solutions. The candidate will ...
This role involves providing technical expertise for Windows Server/Desktop operating systems, administering Active Directory, and managing storage and disaster recovery solutions. The candidate will ...
Azure Databricks Architect
Seattle, WA · On-site
$72.25 - $94.25/hr
The ideal candidate will have deep expertise in Azure Databricks, Spark, Delta Lake, and modern ... Ensure scalability, reliability, security, and disaster recovery across the platform. Nice to Have
Azure Databricks Architect
Seattle, WA · On-site
$72.25 - $94.25/hr
The ideal candidate will have deep expertise in Azure Databricks, Spark, Delta Lake, and modern ... Ensure scalability, reliability, security, and disaster recovery across the platform. Nice to Have
Databricks Engineer
Madison, WI · On-site
... using Spark Structured Streaming (or equivalent), including state management, checkpointing, and recovery. • Design incremental processing, partitioning strategies, and data layout/file sizing ...
Databricks Engineer
Madison, WI · On-site
... using Spark Structured Streaming (or equivalent), including state management, checkpointing, and recovery. • Design incremental processing, partitioning strategies, and data layout/file sizing ...
... recovery, incident / problem / capacity management * Serves as a liaison between client partners ... Experience in Yarn , Spark and Impala job debugging and troubleshooting * Experience in addressing ...
... recovery, incident / problem / capacity management * Serves as a liaison between client partners ... Experience in Yarn , Spark and Impala job debugging and troubleshooting * Experience in addressing ...
Solutions Architect
$71.75 - $94.50/hr
Ensures timely recovery from outages, performs root cause analysis and implements preventative measures. * Automate, document, share, educate, and improve processes. * Develops and manages service ...
Solutions Architect
$71.75 - $94.50/hr
Ensures timely recovery from outages, performs root cause analysis and implements preventative measures. * Automate, document, share, educate, and improve processes. * Develops and manages service ...
Lead DATA ENGINEER - Cloud Migration
Jersey City, NJ · On-site
$107K - $140K/yr
Build scalable batch and streaming pipelines using PySpark, Spark SQL * Leverage Delta Lake for ... Ensure workloads meet operational resilience and recovery expectations * Partner with cloud ...
Lead DATA ENGINEER - Cloud Migration
Jersey City, NJ · On-site
$107K - $140K/yr
Build scalable batch and streaming pipelines using PySpark, Spark SQL * Leverage Delta Lake for ... Ensure workloads meet operational resilience and recovery expectations * Partner with cloud ...
Director of Operations
Minneapolis, MN · On-site
$85K - $100K/yr
... recovery. You ensure the site gets safely back online and the team stays focused. * Build for tomorrow Partner with L&D on Spark Plug and Spark Summit readiness. Strengthen upcoming leaders through ...
Quick apply
Director of Operations
Minneapolis, MN · On-site
$85K - $100K/yr
... recovery. You ensure the site gets safely back online and the team stays focused. * Build for tomorrow Partner with L&D on Spark Plug and Spark Summit readiness. Strengthen upcoming leaders through ...
Director of Operations
Blaine, MN · On-site
$85K - $100K/yr
... recovery. You ensure the site gets safely back online and the team stays focused. * Build for tomorrow Partner with L&D on Spark Plug and Spark Summit readiness. Strengthen upcoming leaders through ...
Director of Operations
Blaine, MN · On-site
$85K - $100K/yr
... recovery. You ensure the site gets safely back online and the team stays focused. * Build for tomorrow Partner with L&D on Spark Plug and Spark Summit readiness. Strengthen upcoming leaders through ...
Working experience L and thorough understanding of SANs, backup systems and disaster recovery, and emergency preparedness strongly desired. Network certifications in Cisco (CCNA or CCNP) and ...
Working experience L and thorough understanding of SANs, backup systems and disaster recovery, and emergency preparedness strongly desired. Network certifications in Cisco (CCNA or CCNP) and ...
Director of Operations
Blaine, MN · On-site
$85K - $100K/yr
... recovery. You ensure the site gets safely back online and the team stays focused. * Build for tomorrow Partner with L&D on Spark Plug and Spark Summit readiness. Strengthen upcoming leaders through ...
Director of Operations
Blaine, MN · On-site
$85K - $100K/yr
... recovery. You ensure the site gets safely back online and the team stays focused. * Build for tomorrow Partner with L&D on Spark Plug and Spark Summit readiness. Strengthen upcoming leaders through ...
Director of Operations
Blaine, MN · On-site
$85K - $100K/yr
... recovery. You ensure the site gets safely back online and the team stays focused. * Build for tomorrow Partner with L&D on Spark Plug and Spark Summit readiness. Strengthen upcoming leaders through ...
Director of Operations
Blaine, MN · On-site
$85K - $100K/yr
... recovery. You ensure the site gets safely back online and the team stays focused. * Build for tomorrow Partner with L&D on Spark Plug and Spark Summit readiness. Strengthen upcoming leaders through ...
... recovery of domain and GPO setup • Fundamental networking concepts and troubleshooting such as DHCP, DNS, Routing, Firewall • Fundamental understanding of security principles, ACL, Encryption • ...
... recovery of domain and GPO setup • Fundamental networking concepts and troubleshooting such as DHCP, DNS, Routing, Firewall • Fundamental understanding of security principles, ACL, Encryption • ...
... SPARK FastTrack Award from Ann Arbor SPARK 2015 -Honoree of Diversity Focused Company by Corp! ... operations Disaster Recovery support Capacity planning Performance tuning 7 plus years of IT ...
... SPARK FastTrack Award from Ann Arbor SPARK 2015 -Honoree of Diversity Focused Company by Corp! ... operations Disaster Recovery support Capacity planning Performance tuning 7 plus years of IT ...
Dataiku Administrator
Austin, TX · On-site
... Hadoop/Spark, Snowflake, etc.). - Configure and manage compute resources, including Spark ... disaster recovery, and business continuity procedures. - Automate administrative tasks and ...
Dataiku Administrator
Austin, TX · On-site
... Hadoop/Spark, Snowflake, etc.). - Configure and manage compute resources, including Spark ... disaster recovery, and business continuity procedures. - Automate administrative tasks and ...
... Hadoop/Spark, Snowflake, etc.). - Configure and manage compute resources, including Spark ... disaster recovery, and business continuity procedures. - Automate administrative tasks and ...
... Hadoop/Spark, Snowflake, etc.). - Configure and manage compute resources, including Spark ... disaster recovery, and business continuity procedures. - Automate administrative tasks and ...
Core Platform Engineer
Alexandria, VA · Hybrid
$65 - $67/hr
Support a highly available and scalable infrastructure containing Object storage, Openshift, Spark ... Failover strategies, disaster recovery, monitoring (Prometheus, Grafana)
Core Platform Engineer
Alexandria, VA · Hybrid
$65 - $67/hr
Support a highly available and scalable infrastructure containing Object storage, Openshift, Spark ... Failover strategies, disaster recovery, monitoring (Prometheus, Grafana)
Dataiku Administrator
Austin, TX · On-site
... Hadoop/Spark, Snowflake, etc.). - Configure and manage compute resources, including Spark ... disaster recovery, and business continuity procedures. - Automate administrative tasks and ...
Dataiku Administrator
Austin, TX · On-site
... Hadoop/Spark, Snowflake, etc.). - Configure and manage compute resources, including Spark ... disaster recovery, and business continuity procedures. - Automate administrative tasks and ...
Only W2- Senior Data Engineer- Onsite in Denver, CO (Only Locals candidates)
Denver, CO · On-site
$117K - $141K/yr
This role focuses on constructing robust, scalable data pipelines using Spark/Scala, ensuring data ... and recovery when upstream issues or schema changes require reprocessing. · Collaborate across ...
Only W2- Senior Data Engineer- Onsite in Denver, CO (Only Locals candidates)
Denver, CO · On-site
$117K - $141K/yr
This role focuses on constructing robust, scalable data pipelines using Spark/Scala, ensuring data ... and recovery when upstream issues or schema changes require reprocessing. · Collaborate across ...
SAP BASIS system admin
Irving, TX · On-site
... SPARK FastTrack Award from Ann Arbor SPARK 2015 -Honoree of Diversity Focused Company by Corp! ... disaster recovery procedures Designing and/or documenting SAP procedures Provide architectural ...
SAP BASIS system admin
Irving, TX · On-site
... SPARK FastTrack Award from Ann Arbor SPARK 2015 -Honoree of Diversity Focused Company by Corp! ... disaster recovery procedures Designing and/or documenting SAP procedures Provide architectural ...
Other
This job post has expired today. Applications are no longer accepted.
Job description
Job Description
We are seeking an experienced Technical Support Engineer at the L3 level to provide onsite support for enterprise data engineering platforms. The successful candidate will serve as the primary technical interface at the customer location, own complex production issues through resolution, and coordinate closely with offshore engineering and support teams. This role requires deep expertise in AWS-based data services, PySpark, AWS Glue, Databricks, Amazon Redshift, and SQL, along with strong debugging, stakeholder-management, and communication skills.
Key Responsibilities
- Provide L3 production support for data engineering applications, ETL/ELT pipelines, batch workflows, and analytical data platforms.
- Own complex incidents from initial triage through resolution, including log analysis, data validation, defect isolation, root cause analysis, recovery, and preventive-action tracking.
- Troubleshoot and resolve issues across AWS Glue, Amazon Redshift, Databricks, PySpark applications, SQL workloads, and related AWS services.
- Diagnose data mismatches, pipeline failures, performance degradation, dependency issues, access problems, and environment or configuration defects.
- Write and optimize complex SQL queries for troubleshooting, reconciliation, data-quality validation, and performance analysis.
- Debug PySpark and Spark-based workloads using execution plans, job logs, cluster metrics, Spark UI, and application diagnostics.
- Monitor production platforms, identify risks proactively, and ensure incidents and service requests are resolved within agreed SLAs.
- Lead incident bridges for critical production issues and provide timely, accurate updates to technical teams, business stakeholders, and customer leadership.
- Coordinate daily with the offshore support and data engineering teams, including work allocation, technical handoffs, follow-ups, knowledge transfer, and status reporting.
- Partner with data engineering, platform, infrastructure, DevOps, security, QA, and vendor teams to implement permanent fixes and maintain platform stability.
- Collaborate with different Business Units at the customer site to understand impact, prioritize issues, clarify requirements, and communicate resolution plans.
- Support release, deployment, change, and post-production validation activities across development, test, and production environments.
- Prepare and maintain runbooks, troubleshooting guides, known-error records, incident reports, RCA documents, and operational dashboards.
- Identify opportunities to automate repetitive support activities, improve monitoring and alerting, and reduce manual recovery effort.
- Mentor L1/L2 support engineers and enable the offshore team to resolve recurring issues independently.
- Participate in an on-call or extended-hours support rotation when required for business-critical incidents.