Senior Data Engineer - Data Pipelines
Own the full lifecycle of large-scale data pipelines, from design to production, across multiple regions. Build and optimize distributed processing jobs, evolve data models, and ensure data quality and observability. Mentor engineers and drive engineering rigor through design reviews, code review, and automation.
Responsibilities
- Design, build, and maintain scalable data pipelines running in multiple regions
- Write and optimize large-scale distributed processing jobs using PySpark on EMR or Databricks
- Evolve data models and pipelines feeding threat intelligence, assets, and detections product layers
- Design and implement data quality frameworks (e.g., Great Expectations) for health checks and conditional monitoring
- Establish and lead observability across all pipelines with distributed tracing, structured logging, and metrics
- Set technical direction for pipeline infrastructure and drive automation of operational tasks
- Conduct design reviews, code reviews, and mentorship to raise overall engineering standards
Requirements
- 7+ years building and operating production data pipelines at scale, including system design and operation
- Deep Python 3.10+ expertise (Pydantic, async, decorators, packaging, profiling)
- Production experience with Apache Airflow at scale: DAG authoring, dynamic task mapping, MWAA or self-hosted
- PySpark mastery: job authoring, optimization (partitioning, shuffle mitigation, join strategies, Catalyst plans), EMR or Databricks
- Snowflake depth: Snowpark, schema/clustering design, query tuning, warehouse sizing, cost attribution
- AWS knowledge: S3, SQS/SNS, EMR, IAM, infrastructure-as-code
- Data quality framework design experience (e.g., Great Expectations) beyond configuration
- Observability tooling: distributed tracing, structured logging, metrics
- CI/CD with artifact packaging and staged deployments across environments and regions
- Strong testing discipline: pytest, moto, integration tests against cloud services