Эта вакансия пока только на английском.

← Назад к вакансиям
M

AI Data Engineer

Millennium·Израиль·en
В офисеПолная занятостьData EngineeringInvestment Management

Experienced data engineer focused on building scalable ETL and document processing pipelines for AI and LLM applications. Skilled in Python, OCR, chunking, and retrieval systems to ensure high-quality data for RAG.

Обязанности

  • Design, build, and maintain scalable ETL and data ingestion pipelines from file shares, object stores, APIs, and databases.
  • Develop document understanding workflows including parsing, layout analysis, OCR, text and metadata extraction, and normalization for PDF, Office, HTML, and images.
  • Implement chunking, cleaning, and enrichment strategies to improve retrieval quality for downstream RAG systems.
  • Build change-detection, deduplication, and incremental update mechanisms for large document corpora.
  • Engineer pipelines for correctness, throughput, and resilience, handling malformed inputs, large files, and high volumes.
  • Establish data quality checks, observability, and metrics for early issue identification and fast resolution.
  • Partner with stakeholders to understand source systems and content requirements, translating them into production-ready ingestion solutions.
  • Stay current with AI, LLM, document AI, and retrieval advances, applying improvements to team solutions.

Требования

  • 4+ years of experience with strong proficiency in Python, including building data pipelines, services, and APIs.
  • Hands-on experience designing and developing ETL and data pipeline solutions for large data volumes.
  • Experience with document processing and text extraction, including PDF and Office parsing, OCR, and unstructured content.
  • Solid understanding of data modeling, transformation, and data quality best practices.
  • Experience designing, building, testing, and debugging high-performance, reliable systems.
  • Clear communication skills, able to explain complex technical concepts to both technical and non-technical audiences.

Будет плюсом

  • Familiarity with RAG systems and the impact of ingestion on retrieval quality, including chunking strategies, embeddings, and vector stores.

Соответствие

Больше возможностей

Похожие вакансии

Новые вакансии в категории «Data Engineering».

Создайте профиль, чтобы увидеть оценку соответствия.

Создайте профиль, чтобы увидеть оценку соответствия.

Создайте профиль, чтобы увидеть оценку соответствия.