Senior Data Engineer
Join a fast-growing startup transforming infrastructure monitoring with optical fibers and AI. Build scalable data pipelines, optimize storage, and collaborate with cross-functional teams to enable real-time analytics and AI model training.
Responsibilities
- Design, develop, and maintain scalable data pipelines integrating data from APIs, databases, streaming platforms, edge devices, cloud services, and on-premise systems.
- Build reliable batch and real-time data processing workflows for business-critical use cases.
- Develop data access services and APIs for efficient communication between edge devices, cloud, and on-premise environments.
- Design and optimize storage solutions for structured, semi-structured, and high-dimensional sensor data.
- Implement scalable streaming data pipelines for low-latency processing and event-driven architectures.
- Analyze existing data architecture and workflows to improve performance, scalability, reliability, and maintainability.
- Optimize cloud infrastructure and storage costs while maintaining high availability and low latency.
- Collaborate with Software, AI, Algorithms, DevOps, and Product teams to translate requirements into technical solutions.
- Design monitoring, observability, and operational processes for data platforms.
Requirements
- Bachelor's degree in Computer Science, Engineering, Mathematics, or a related quantitative field.
- 5+ years of professional experience in Data Engineering or a related role.
- Strong experience designing large-scale data pipelines using orchestration frameworks such as Apache Airflow or Prefect.
- 5+ years of software development experience, including at least 2 years of Python development.
- Strong knowledge of relational and NoSQL databases such as PostgreSQL, MySQL, MongoDB, Elasticsearch, ClickHouse, or similar.
- Experience designing streaming and event-driven architectures with technologies like AWS Kinesis, SQS, RabbitMQ, Kafka, or similar.
- Experience designing REST APIs and backend services using FastAPI or similar frameworks.
- Experience with AWS cloud services (S3, EC2, Lambda, CloudWatch, IAM, etc.).
- Experience with Git, Docker, CI/CD pipelines, and modern software engineering practices.
- Excellent communication and collaboration skills with engineering, AI, and Product teams.
- Self-driven, innovative, and continuously looking for ways to improve systems and processes.
Nice to have
- Experience with Kubernetes and container orchestration.
- Experience with distributed computing platforms and distributed data processing systems.
- Experience building ML data pipelines supporting training and inference workloads.
- Experience working with large-scale sensor, IoT, or time-series data.
- Experience with monitoring and observability tools such as Grafana, Prometheus, ELK, Kibana, or OpenSearch.
- Experience working in edge computing or hybrid cloud environments.