Must be Local only within 40 Miles area
Relocation will not work
Look insurance domain highly preferred
Job Description:
11 Years exp
We are looking for a Senior Data Engineer to design, build, and optimize scalable data pipelines and warehouse solutions that power business-critical analytics. You will work closely with data architects, analysts, and business stakeholders to deliver robust data products on modern lakehouse platforms.
Key Responsibilities
- Design and develop end-to-end data pipelines using Apache Spark and Databricks, following medallion (Bronze/Silver/Gold) architecture patterns
- Build and maintain large-scale SQL-based data warehouses, including dimensional models, star/Client schemas, and performance-tuned queries
- Lead data ingestion from diverse sources (RDBMS, APIs, flat files, streaming) into centralized platforms with strong data quality controls
- Implement and enforce Unity Catalog governance standards — data lineage, access controls, tagging, and cataloging
- Optimize Spark jobs for performance, cost efficiency, and reliability at scale
- Collaborate with architects to define standards for data modeling, pipeline design, and naming conventions
- Mentor junior engineers and conduct code reviews to uphold engineering best practices
- Partner with business analysts and data consumers to translate requirements into scalable data solutions
- Proactively identify and resolve data quality, latency, and pipeline reliability issues
Required Skills & Qualifications
- 10+ years of hands-on experience in data engineering
- Strong expertise in Apache Spark (PySpark / Scala) and Databricks platform
- Deep proficiency in SQL — query optimization, window functions, complex transformations, stored procedures
- Solid experience with data warehousing concepts — normalization, SCD types, fact/dimension modeling
- Experience with Client Lake or similar open table formats (Apache Iceberg, Hudi)
- Hands-on with orchestration tools such as Apache Airflow, Databricks Workflows, or Azure Data Factory
- Familiarity with version control (Git) and CI/CD practices for data pipelines
- Strong understanding of data governance — lineage, cataloging, data quality frameworks
- Excellent problem-solving skills and ability to work independently in a client-facing environment
Good to Have
- Experience with dbt (data build tool) for transformation layer development
- Exposure to cloud platforms — AWS
- Knowledge of streaming technologies (Kafka, Event Hubs)
- Familiarity with Great Expectations or other data quality frameworks