Hi,
ย
Please find the job description below
ย
Role: Data Bricks Engineer
Location: Plano, TX
ย
Job Summary:
ย We are seeking an experienced Databricks Engineer with strong expertise in migration of PySpark/Hive workloads from AWS EMR or legacy platforms data platforms while ensuring data quality, security, and scalability
ย
Key Responsibilities:
ย
โขโโโโโDesign, develop, and maintain scalable data pipelines using Databricks Implement and manage Unity Catalog for centralized data governanance
โขโโโโโDevelop and optimize Spark Declarative Pipelines (SDP) and moderatePerform end-to-end data validation, reconciliation, and quality assurance.
โขโโโโโLead migration projects involving:
โขโโโโโPySpark jobs from AWS EMR to Databricks Hive-based ETL workloads to Databricks and Delta Lake.
โขโโโโโLegacy data platforms to Databricks Lakehouse architecture.
โขโโโโโConvert and optimize Hive SQL, Spark SQL, and PySpark workloads forImple and automation for validation and PySparkjobs from AWS EMR to Databricks.
โขโโโโโHive-based ETL workloads to Databricks and Delta Lake.
โขโโโโโLegacy data platforms to Databricks Lakehouse architecture.
โขโโโโโConvert and optimize Hive SQL, Spark SQL, and PySpark workloads for Database Implement data quality frameworks and automation for validation and reconcilation Work closely with architects, business stakeholders, and data consumers to | Optimize Databricks jobs for performance, reliability, and cost efficiency.
โขโโโโโManage CI/CD processes and infrastructure automation for Databricks deplo Ensure compliance with enterprise data security, governance.
ย
Required Skills:
ย
โขโโโโโDatabricks & Data Engineering Strong experience with Databricks Lakehouse Platform.
โขโโโโโHands-on experience with Delta Lake, Medallion Architecture, and Databric Expertise in PySpark and Spark SQL.
โขโโโโโExperience working with large-scale data processing and performance tuning Unity Catalog Strong understanding of: Unity Catalog setup and administration.
โขโโโโโStrong understanding of Unity Catalog setup and administration Fine-grained access control (RBAC). Data lineage and audit capabilities.
โขโโโโโData governance and security implementation.
โขโโโโโSpark Declarative Pipelines (SDP) Hands-on experience implementing and maintaining SDP pipelines.
โขโโโโโKnowledge of pipeline orchestration, monitoring, and optimization.
โขโโโโโExperience building reusable and scalable data transformation framework Data Validation & Quality Experience designing and implementing data validation frameworks.
โขโโโโโData reconciliation between source and target systems. Validation of migrated datasets to ensure accuracy and completeness.
โขโโโโโExperience with data quality monitoring and observability tools.
โขโโโโโMigration Experience Proven experience in migratingPySparkworkloads from AWS EMR to Databricks.
โขโโโโโMive-based ETL jobs to Databricks. Legacy data warehouses and Hadoop ecosystems to Databricks.
ย
ย