Job Description
We are seeking a Databricks ETL Developer with hands-on experience in building scalable ETL pipelines using the Databricks platform. The ideal candidate should have strong expertise in Spark, SQL, Python, cloud data engineering, and modern data warehousing concepts. A Databricks Certification is highly preferred.
Required Skills
- 5+ years of experience in Data Engineering and ETL development.
- Hands-on experience with Databricks and Apache Spark (PySpark).
- Strong experience designing and developing ETL/ELT pipelines.
- Expertise in Python, SQL, and Spark SQL.
- Experience with Delta Lake, Unity Catalog, and Delta Live Tables (DLT).
- Experience with Azure Databricks, AWS Databricks, or Databricks on Google Cloud Platform.
- Knowledge of Data Lake, Data Warehouse, and Medallion Architecture.
- Experience integrating data from relational databases, APIs, flat files, and cloud storage.
- Strong understanding of data modeling, partitioning, optimization, and performance tuning.
- Experience with orchestration tools such as Azure Data Factory (ADF), Apache Airflow, or similar.
- Familiarity with Git, CI/CD pipelines, and Agile methodologies.
- Excellent troubleshooting and analytical skills.
Responsibilities
- Design, develop, and maintain scalable ETL/ELT pipelines using Databricks.
- Process large-scale structured and unstructured datasets using PySpark.
- Develop optimized Spark jobs for batch and incremental data processing.
- Build and maintain Delta Lake tables for reliable data storage.
- Perform data validation, cleansing, and transformation activities.
- Optimize Databricks workloads for performance and cost efficiency.
- Collaborate with Data Architects, Analysts, and Business teams to deliver data solutions.
- Implement data quality checks and monitoring processes.
- Support production deployments and resolve data pipeline issues.
- Document ETL processes and technical designs.
Preferred Qualifications
- Databricks Certified Data Engineer Associate or Databricks Certified Data Engineer Professional.
- Experience with Azure Data Factory (ADF), Snowflake, Azure Synapse, or Microsoft Fabric.
- Experience with Kafka, Event Hubs, or other streaming technologies.
- Knowledge of Power BI, Tableau, or other BI tools.
- Exposure to DevOps and Infrastructure as Code (Terraform) is a plus.