Senior Data Engineer
Own the full lifecycle of production data pipelines across more than 15 regions, supporting threat intelligence, asset, and detection data for a cybersecurity platform.
Responsibilities
- Build and maintain data pipelines across more than 15 regions.
- Develop distributed processing jobs for large datasets.
- Create pipelines that supply access-layer data for threat intelligence, assets, and detections.
- Build data health frameworks with conditional checks and observability.
- Implement distributed tracing across data pipelines.
- Support production monitoring, security scanning, and on-call operations.
Requirements
- At least 4 years of experience building production data pipelines at scale.
- Strong Python 3.10+ skills, including Pydantic, asynchronous programming, and decorators.
- Production Apache Airflow experience, including DAG authoring and MWAA or self-hosted deployments.
- Experience authoring and optimizing PySpark jobs with EMR or Databricks.
- Expertise with Snowflake, including Snowpark and schema design.
- Knowledge of AWS services including S3, SQS/SNS, and EMR.
- Experience with data quality tools such as Great Expectations or Monte Carlo.
- Experience with CI/CD, artifact packaging, and staged deployments.
- Strong testing practices using pytest, moto, and integration tests with cloud services.