Role: Data Engineer
Location: Mount Laurel, NJ (3 Days onsite/week)
Contract
Job Description:
We are looking for a skilled Data Engineer with strong hands-on experience in Azure Databricks and Azure Data Factory to design, develop, and maintain scalable data pipelines and cloud-based data solutions. The ideal candidate will work closely with data architects, analysts, business stakeholders, and application teams to build reliable ETL/ELT workflows, transform large datasets, and enable analytics and reporting use cases.
Key Responsibilities:
• Design, develop, and maintain scalable data pipelines using Azure Data Factory and Azure Databricks.
• Build ETL/ELT workflows for data ingestion, transformation, cleansing, enrichment, and loading into data lakes or data warehouses.
• Develop and optimize PySpark, Spark SQL, and SQL scripts for large-scale data processing.
• Integrate data from multiple sources such as databases, APIs, flat files, cloud storage, and enterprise applications.
• Implement incremental data loading, scheduling, parameterization, and reusable pipeline frameworks in ADF.
• Work with Azure Data Lake Storage, Delta Lake, and lakehouse architecture patterns for structured and unstructured data.
• Monitor, troubleshoot, and optimize data pipelines for performance, reliability, cost efficiency, and data quality.
• Collaborate with business users, data analysts, data scientists, and architects to understand requirements and deliver data solutions.
• Implement data validation, reconciliation, exception handling, logging, and alerting mechanisms.
• Follow coding standards, version control, CI/CD, and deployment best practices using tools such as Git and Azure DevOps.
• Ensure data security, access control, and compliance by applying Azure security best practices.
• Prepare technical documentation including data flow diagrams, mapping documents, pipeline design, and support guides.
Required Technical Skills:
• Strong hands-on experience with Azure Databricks, Databricks notebooks, clusters, jobs, and workflows.
• Strong experience in Azure Data Factory including pipelines, datasets, linked services, triggers, parameters, variables, and integration runtimes.
• Good programming experience in PySpark, Python, Spark SQL, and SQL.
• Experience working with Azure Data Lake Storage Gen2, Delta Lake, and data lakehouse concepts.
• Knowledge of ETL/ELT design patterns, data warehousing concepts, dimensional modeling, and data integration methods.
• Experience in performance tuning of Spark jobs, SQL queries, and ADF pipelines.
• Understanding of batch processing, incremental loading, CDC concepts, and data partitioning strategies.
• Experience with Git, Azure DevOps, CI/CD pipelines, and release management for data engineering solutions.
• Knowledge of data quality checks, monitoring, logging, and error handling frameworks.
• Basic understanding of Azure security concepts such as managed identities, service principals, Key Vault, RBAC, and private endpoints.
Preferred / Good-to-Have Skills:
• Experience with Azure Synapse Analytics, Azure SQL Database, SQL Server, or Snowflake.
• Knowledge of Unity Catalog, data governance, metadata management, and data lineage.
• Experience with streaming data processing using Kafka, Event Hubs, or Databricks Structured Streaming.
• Exposure to Power BI, reporting platforms, or analytics consumption layers.
• Experience working in Agile delivery models and cross-functional project teams.
• Microsoft Azure Data Engineer certification or Azure Databricks certification is an added advantage.