About the Role
We're looking for an experienced data engineer to design, build, and operate our real-time event-processing pipeline. You'll own a system that ingests high-volume event data (100M+ events/day) from multiple producers, processes it through Apache Flink for enrichment, deduplication, and aggregation, and delivers results to downstream services and APIs with low latency and strong correctness guarantees.
Responsibilities
- Design and maintain Kafka-based ingestion pipelines, including topic strategy, partitioning, and producer/consumer contracts
- Build and optimize Apache Flink streaming jobs (Java or Scala) for real-time transformation, deduplication, and event correlation
- Implement event-time processing using watermarks to correctly handle out-of-order and late-arriving events
- Ensure exactly-once processing semantics and manage Flink checkpointing/state backends for fault tolerance
- Diagnose and resolve production issues: consumer lag, checkpoint failures, back pressure, schema mismatches
- Tune Flink jobs for latency, throughput, and resource efficiency
- Collaborate on system architecture from ingestion through storage to API exposure
- Participate in on-call rotation and incident root-cause analysis
Requirements
- Strong experience with Apache Flink in production (Java or Scala)
- Solid understanding of Apache Kafka: producers, consumers, partitioning, offset management
- Experience with event-time semantics, watermarks, and stateful stream processing
- Familiarity with checkpointing, state backends (RocksDB, etc.), and recovery strategies
- Experience debugging distributed systems issues (back pressure, lag, failures) using logs/metrics
- Understanding of exactly-once vs at-least-once delivery guarantees
- Experience designing systems handling 100M+ daily events is a strong plus
- Comfortable with live coding/system design during interviews
Nice to Have
- Migration experience from Spark to Flink
- Experience exposing streaming results via REST APIs
- Familiarity with storage systems for high-throughput write patterns (e.g., Cassandra, Druid, ClickHouse)