AI Infrastructure Engineer
Build and scale the compute, data, and deployment infrastructure supporting AI training and inference workloads. The role covers GPU clusters, ML pipelines, production operations, and collaboration with research and product engineering teams. Hybrid work is based in Tel Aviv.
Responsabilidades
- Design, build, and operate infrastructure for AI training and inference workloads
- Manage and optimize GPU clusters for performance, utilization, and cost
- Build and maintain data and model pipelines, experiment tracking, and reproducible training environments
- Own CI/CD, orchestration, and observability for production ML systems
- Collaborate with research and product engineers to move models from prototypes into production
Requisitos
- At least 4 years of experience in infrastructure, DevOps, platform, or ML engineering
- Strong Python and Linux skills with solid software engineering fundamentals
- Hands-on experience with Kubernetes, Docker, and infrastructure as code such as Terraform
- Experience with AWS, GCP, or Azure and GPU-based workloads
- Familiarity with PyTorch, Ray, Slurm, or MLflow
Se valora
- Experience with distributed training at scale
- Background in EDA, chip design, or other HPC-intensive domains
- Experience with CUDA, NCCL, or low-level performance optimization
Beneficios
- Work on deep-tech problems at the intersection of AI and silicon
- Join a small, senior engineering team with direct product impact
- Hybrid work from a Tel Aviv office