Data Engineer
Hiring a Data Engineer to design and maintain scalable, low-latency data pipelines and big-data infrastructure for high-throughput real-time systems. You will own data modeling, optimization, observability, and collaborate across engineering, security, and data science teams.
Responsibilities
- Design, build, and maintain scalable, low-latency data pipelines, ETL/ELT, and infrastructure.
- Architect data models, schemas, and storage layouts optimized for fast retrieval at large volume.
- Build real-time and event-driven ingestion and processing for high-throughput data.
- Optimize SQL and NoSQL databases, including query tuning, indexing, and partitioning.
- Work with big-data volumes across data lakes, warehouses, and lakehouses using Spark.
- Own data observability, including pipeline health, freshness, and quality.
- Collaborate with software engineers, security experts, and data scientists to integrate data into products.
Requirements
- B.Sc. in Computer Science, Data Engineering, or a related field.
- 3+ years of hands-on experience with large-scale data infrastructure.
- Strong Python skills, including Pandas.
- Deep SQL and NoSQL knowledge with performance-optimization experience.
- Proven experience with real-time/low-latency systems and data architecture.
- Comfort with big-data volumes, pipelines, and data lakes.
- Experience with streaming/event-driven design, such as Kafka.
- Spark/PySpark experience for large-scale processing.
- Strong problem-solving skills and a proactive, independent mindset.
Nice to have
- Experience with ClickHouse.
- Data warehousing and lakehouse/open table formats, e.g., Apache Iceberg.
- Vector databases and embeddings, e.g., pgvector.
- Strong system design and data architecture sense.
- Experience with AI-assisted development tools, e.g., Cursor, Claude Code.
- Experience with security, fraud, or high-volume telemetry data.
- Containerized infrastructure, e.g., Docker, Kubernetes, and DevOps basics.
- ElasticSearch, Redis, DynamoDB, or Metabase.