Mid-Level Generative AI Engineer
Develop, deploy, and maintain generative AI and LLM solutions, including RAG systems, AI agents, enterprise search, and local models in secure on-premise or air-gapped environments.
Responsabilidades
- Develop and implement AI and LLM solutions in on-premise environments.
- Deploy and operate GPU-based local AI models.
- Develop RAG solutions, AI agents, and enterprise search engines.
- Install and deploy models using Ollama, vLLM, llama.cpp, Hugging Face, or similar tools.
- Use AI-powered development tools for code generation, analysis, refactoring, testing, documentation, and process automation.
- Use collaborative AI tools for knowledge sharing, document analysis, and engineering support.
- Optimize models through prompt engineering, tool calling, quantization, and fine-tuning.
- Onboard, scan, transfer, and deploy models and packages into isolated environments.
- Monitor response times, accuracy, CPU/GPU utilization, memory consumption, and computational costs.
Requisitos
- Proven experience with Python software development.
- Experience developing backend services and REST APIs.
- Hands-on experience with LLMs and generative AI solutions.
- Experience deploying and operating local models with Ollama, vLLM, llama.cpp, or similar tools.
- Experience with Hugging Face and open-source models.
- Experience developing RAG solutions, embeddings, and vector databases.
- Experience with Linux and Docker.
- Familiarity with GPU servers, CUDA, and compute resource management.
- Experience integrating AI solutions with enterprise systems and internal data sources.
- Experience with Git and CI/CD processes.
- Experience using Claude Code or similar AI-powered development tools.
- Ability to work in secure, communication-restricted, or air-gapped environments.