Principal MLOps Engineer
Design and scale ML and LLMOps platforms that enable data scientists and security researchers to train, deploy, serve, and continuously improve advanced AI systems.
Обязанности
- Design and optimize distributed GPU infrastructure for LLM and SLM training and fine-tuning
- Architect automated continuous training and deployment pipelines across the ML lifecycle
- Own production model-serving architecture, balancing latency, throughput, and GPU utilization
- Build observability systems for model performance, data drift, and compute metrics
- Feed monitoring insights into automated training and continuous-improvement workflows
- Partner with data scientists and security researchers to productionize complex model architectures
- Integrate ML platforms with core cloud infrastructure in collaboration with DevOps teams
Требования
- 4+ years of hands-on experience as a Senior ML Engineer, MLOps Engineer, or Backend Platform Engineer in cloud environments
- Experience managing the technical lifecycle of classic ML, LLM/SLM, and agentic or RAG systems
- Expert Python skills for ML infrastructure, pipelines, and automation
- Experience designing scalable data preparation and processing pipelines
- Experience with distributed multi-GPU training using tools such as PyTorch, DeepSpeed, Megatron-LM, or cloud-native training infrastructure
- Strong knowledge of deep learning concepts, training dynamics, and optimization techniques
- Infrastructure expertise with AWS, Azure, or GCP managed AI platforms and services
- Experience integrating CI/CD patterns such as GitLab CI or GitHub Actions into software and model delivery
- Proficiency with AI development tools and ecosystems for generating, reviewing, and testing code
- Applicants must be able to work in Israel; immigration sponsorship is not available
Будет плюсом
- Strong experience with the GCP ecosystem
- Background in data science or deep learning workflows
- Cybersecurity domain knowledge