Senior Data Engineer
Join a team building a cloud data platform processing millions of records per second. You will own the analytical storage layer, design batch processing pipelines using Spark, and develop high-performance APIs. This role focuses on large-scale data modeling, query optimization, and storage efficiency at petabyte scale.
Responsibilities
- End-to-end ownership of large-scale analytical storage layer, including data modeling, schema design, partitioning, and query optimization
- Design and develop batch processing layer over data lake using Spark on EMR (Java, PySpark)
- Build and evolve Java/Spring Boot services exposing data through APIs
- Own performance and cost optimization: query tuning, file layout, cluster sizing, and storage efficiency
- Research and adopt new lakehouse and analytical storage technologies
Requirements
- 5+ years designing and developing large-scale distributed data systems
- Deep expertise in at least one of: columnar/analytical databases (ClickHouse preferred), Apache Spark at scale, or open table formats (Iceberg, Delta Lake, Hudi)
- Experience with data lake technologies: Parquet, S3, SQL query engines (Athena, Trino, Presto)
- Strong command of analytical data modeling and design principles
- Strong Java and object-oriented design skills
- Experience building microservices on Kubernetes
- Hands-on experience with AWS (EMR, S3, Glue)
- B.Sc. in Computer Science or related field, or equivalent experience