← Back to jobs
M

AI Data Engineer

·Israel
On-siteFull-timeData EngineeringInvestment Management

Experienced data engineer focused on building scalable ETL and document processing pipelines for AI and LLM applications. Skilled in Python, OCR, chunking, and retrieval systems to ensure high-quality data for RAG.

Responsibilities

  • Design, build, and maintain scalable ETL and data ingestion pipelines from file shares, object stores, APIs, and databases.
  • Develop document understanding workflows including parsing, layout analysis, OCR, text and metadata extraction, and normalization for PDF, Office, HTML, and images.
  • Implement chunking, cleaning, and enrichment strategies to improve retrieval quality for downstream RAG systems.
  • Build change-detection, deduplication, and incremental update mechanisms for large document corpora.
  • Engineer pipelines for correctness, throughput, and resilience, handling malformed inputs, large files, and high volumes.
  • Establish data quality checks, observability, and metrics for early issue identification and fast resolution.
  • Partner with stakeholders to understand source systems and content requirements, translating them into production-ready ingestion solutions.
  • Stay current with AI, LLM, document AI, and retrieval advances, applying improvements to team solutions.

Requirements

  • 4+ years of experience with strong proficiency in Python, including building data pipelines, services, and APIs.
  • Hands-on experience designing and developing ETL and data pipeline solutions for large data volumes.
  • Experience with document processing and text extraction, including PDF and Office parsing, OCR, and unstructured content.
  • Solid understanding of data modeling, transformation, and data quality best practices.
  • Experience designing, building, testing, and debugging high-performance, reliable systems.
  • Clear communication skills, able to explain complex technical concepts to both technical and non-technical audiences.

Nice to have

  • Familiarity with RAG systems and the impact of ingestion on retrieval quality, including chunking strategies, embeddings, and vector stores.

Relevance

More opportunities

Similar jobs

Finding the best alternatives for you…

Questions, answered

Frequently asked questions