Title: Data Platform Engineer
Location: Austin, TX or Sunnyvale, CA (Hybrid)
Duration: 6 months (possibility of extension)
Implementation Partner: Infosys
End Client: To be disclosed
JD:
Job Summary
We are seeking an experienced Data Platform Engineer to architect, build, and operate our cloud-based data lakehouse platform. The role is responsible for developing scalable
data infrastructure and processing pipelines built on open table formats, distributed query and compute engines, and containerized cloud infrastructure - enabling analysts, data
scientists, and downstream applications to access reliable, high-performance data at scale. The ideal candidate combines deep distributed-systems and data engineering expertise with
strong software engineering practices and a passion for building robust, efficient platforms.
Key Responsibilities
Lakehouse Architecture (Apache Iceberg)
- Architect, implement, and maintain the data lakehouse using Apache Iceberg table formats.
- Manage table lifecycle including partitioning, schema evolution, snapshot and metadata management, and compaction.
- Optimize table layout and file sizing for query performance, storage efficiency, and cost.
- Implement data governance, time-travel, and ACID transaction patterns across the lakehouse.
Distributed Query & Compute (Trino & Spark)
- Deploy, tune, and operate Trino clusters for interactive and federated analytics; manage connectors, catalogs, and resource groups.
- Build and optimize large-scale batch and streaming data processing with Apache Spark.
- Troubleshoot query plans, job performance, spills, and shuffle bottlenecks across distributed workloads.
- Balance capacity, concurrency, and cost across query and compute engines.
Data Engineering (Scala & Python)
- Develop scalable data pipelines, frameworks, and services in Scala and Python.
- Build reusable ingestion, transformation, and data-quality libraries for batch and streaming.
- Write clean, well-tested, performant code following software engineering best practices.
- Integrate pipelines with the Iceberg lakehouse and Trino/Spark processing layers.
Cloud & Kubernetes Infrastructure
- Design and operate data platform components on Kubernetes (EKS, AKS, GKE, or self-managed).
- Build and manage containerized workloads, Helm charts, and autoscaling for data services.
- Provision and manage cloud infrastructure (AWS, Azure, or GCP), including object storage and compute.
- Implement infrastructure as code, CI/CD, monitoring, and observability for the platform.
Platform Reliability & Collaboration
- Ensure the reliability, scalability, and security of the data platform across environments.
- Automate deployment, testing, and operational workflows to eliminate toil.
- Partner with data consumers to understand requirements and improve data accessibility.
- Mentor engineers and promote best practices in data and platform engineering.