Data Engineer
Design, develop, and optimize data pipelines and workflows. Responsibilities include complex SQL, data warehouse architecture, and data quality monitoring. Requires strong experience with Apache Spark, Kafka, Apache Iceberg or ClickHouse/Trino, Java, AWS, and Python. Nice to have: Databricks, Flink, Avro, Delta Lake. Full-time role.
Responsibilities
- Design, develop, and optimize data pipelines and workflows
- Write and maintain complex SQL queries including transformations, window functions, and stored procedures
- Design and maintain data warehouse schemas and architecture
- Monitor data quality, integrity, and pipeline performance
- Collaborate with cross-functional teams to ensure clean data integration
- Troubleshoot and resolve data issues
Requirements
- Apache Spark (Structured Streaming and batch) on EMR or equivalent
- Apache Kafka, Apache Iceberg, or ClickHouse/Trino (strong experience in at least one, familiarity with the others)
- Java for ingestion services and Kafka producers/consumers
- Strong SQL skills: window functions, MERGE patterns, complex transformations
- AWS: MSK, EMR, ECS, S3, Glue, MWAA, EventBridge
- Python: Airflow DAGs, data quality, scripting
- Experience with AI coding assistants as a productivity tool
Nice to have
- ClickHouse or Trino/Starburst
- Databricks
- Flink
- Avro, schema registries
- Delta Lake or Hudi