Job Title: DDCS Data Engineer
Key Skills: Delivery, Devices, and Connected Solutions (DDCS), Microsoft Azure Fabric (Lakehouse, Data Factory, Fabric Pipelines), Delta Lake, Azure Data Factory, Azure Blob Storage, Azure Functions, Azure AI Services, Azure RBAC, Managed Identities, Key Vault.
Experience: 8-10 Years’ experience
Location: Indiana Polis, Indiana
We at Coforge are hiring experienced professionals with strong experience in Azure Fabric-based data pipelines, API integrations, AI/RAG document ingestion solutions, and PostgreSQL data platforms to support Lilly's regulated data and AI ecosystem.
Key Responsibilities:
- Design, build, and maintain data ingestion pipelines using Microsoft Azure Fabric Lakehouse and PostgreSQL.
- Develop ETL/ELT pipelines with data validation, quality checks, audit trails, and controlled data governance.
- Build MCP connectors and AI-powered document ingestion pipelines for SharePoint, OneDrive, Veeva, Jama, and other document sources.
- Implement RAG (Retrieval-Augmented Generation) solutions using document chunking, embeddings, vector databases, and LLMs.
- Develop REST API integrations with OAuth authentication, pagination, schema management, and error handling.
- Create OCR-based document processing pipelines for scanned and handwritten PDFs using Azure AI Document Intelligence or similar tools.
- Build and manage data solutions on Azure Fabric, including Fabric Data Factory, Lakehouse, Delta Lake, and Fabric Pipelines.
- Develop and support AWS-based data integrations using S3, Lambda, API Gateway, Glue, and RDS/Aurora.
- Design and implement PostgreSQL Gold-layer schemas, data models, lineage tracking, and governance frameworks.
- Develop production-grade solutions using Python, PySpark, Docker, CI/CD, and Git.
- Implement data quality monitoring, operational alerting, and pipeline documentation.
- Integrate Azure OpenAI/LLM APIs into enterprise AI search and knowledge retrieval solutions.
- Ensure compliance with GxP, ALCOA+, 21 CFR Part 11, and enterprise security standards.
- Collaborate with architects and stakeholders on schema evolution, data modeling, and platform scalability.
Required Skills & Experience
- Microsoft Azure Fabric (Lakehouse, Data Factory, Delta Lake), Python / PySpark, ETL/ELT Pipeline Development, PostgreSQL & SQL, REST APIs & OAuth, Azure OpenAI, RAG & Vector Search.
- Document Ingestion, OCR & AI Search, Azure AI Services & Azure Data Factory, AWS (S3, Lambda, API Gateway, Glue), Data Modeling & Medallion Architecture.
- Docker, CI/CD & Git, MCP Connectors, Data Lineage, Audit Trail & Data Governance, ALCOA+, GxP, 21 CFR Part 11 Compliance, Production-Grade Data Engineering Experience (5+ Years).