Role : Data Engineer/Python Developer
Location:- Alanta, GA (Client in Person interview)
Job Type: Fulltime
Primary Skills : Azure Data Bricks, Data Factory, Pyspark and Master Data Management
Job Description :
A seasoned Data Engineer specialising in enterprise-scale data platform design across Databricks and Microsoft Fabric (Azure), with a technology-agnostic philosophy that delivers portable, future-proof solutions. Recognised for deep expertise in Medallion architecture, metadata-driven pipeline orchestration, and distributed processing with Apache Spark, paired with a strong command of data governance, Master Data Management, and enterprise catalog tooling. Rounds out a comprehensive engineering profile with disciplined CI/CD practices, Infrastructure-as-Code, and data quality observability frameworks that ensure reliable, production-grade data systems at scale.
โขโโโโโProficient across Databricks and Microsoft Fabric (Azure) with a technology-agnostic approach, designing portable solutions that leverage the strengths of each platform interchangeably.
โขโโโโโArchitects and implements Medallion (Bronze/Silver/Gold) data lake frameworks, enforcing clear separation of raw ingestion, conformance, and curated analytical layers.
โขโโโโโDesigns and orchestrates fault-tolerant, scalable data pipelines using Azure Data Factory, Databricks Workflows, and Microsoft Fabric Pipelines โ augmented by metadata-driven, configuration-as-code automation frameworks that enable dynamic pipeline generation, parameterization, and self-service onboarding of new data sources with minimal manual effort.
โขโโโโโApplies Apache Spark (PySpark / Spark SQL) for large-scale distributed data processing, transformation, and performance-tuned query optimization across batch and streaming workloads.
โขโโโโโImplements Master Data Management (MDM) solutions including golden-record creation, entity resolution, probabilistic/deterministic matching, and deduplication to ensure a single source of truth.
โขโโโโโEnforces data governance, stewardship, and data-contract standards โ defining ownership, access policies, SLA commitments, and end-to-end lineage โ while configuring enterprise data catalogs via Unity Catalog (Databricks) and Microsoft Purview (Fabric/Azure) for asset discovery, classification, sensitivity labeling, and access control.
โขโโโโโEstablishes data quality frameworks and observability pipelines with automated profiling, anomaly detection, and SLA monitoring to proactively detect and remediate data issues in production.
โขโโโโโApplies CI/CD practices, Git-based version control, Infrastructure-as-Code (IaC), and rigorous unit/integration testing for Python and SQL codebases to ensure reliable, repeatable deployments.