Senior Applied AI Evaluation Engineer
A telecommunications technology company is seeking a hands-on engineer to evaluate and improve AI systems used for RAN engineering ticket intelligence, root-cause analysis, ownership recommendations, and workflow improvement.
Responsabilidades
- Define RAN ticket-intelligence use cases, acceptance criteria, evaluation metrics, and quality guardrails.
- Build and maintain versioned evaluation datasets from resolved tickets, root-cause analyses, logs, test evidence, code changes, reviews, reassignment history, and outcomes.
- Establish current-tool and non-AI baselines before evaluating LLM, RAG, search, or agent approaches.
- Evaluate internal, commercial, local, and open-source solutions using secure and reproducible data-handling processes.
- Measure retrieval quality, groundedness, diagnosis accuracy, citation support, routing recommendations, calibration, abstention, latency, cost, and human effort.
- Design held-out, time-based, edge, and adversarial test cases while preventing data leakage and future-outcome contamination.
- Analyze failures and prioritize improvements for incorrect conclusions, misrouting, unsupported claims, and missed evidence.
- Prototype improvements in search, metadata, context construction, prompting, reranking, classification, agent workflows, and model selection.
- Develop reusable evaluation pipelines, tools, services, APIs, dashboards, and documented workflows.
- Collaborate with AI, RAN, QA, system integration, release, field, data, and engineering teams to review results and support evidence-based decisions.
- Analyze engineering workflows to identify bottlenecks, handoffs, dependencies, and process-improvement opportunities.
Requisitos
- Bachelor’s or master’s degree in computer science, data science, machine learning, statistics, electrical engineering, or a related field, or equivalent practical experience.
- At least five years of hands-on experience in applied machine learning, data science, search, natural-language processing, analytics engineering, or AI-enabled software systems.
- Recent experience evaluating LLM, RAG, search, or agent systems with representative datasets, task-specific metrics, human review, failure analysis, and regression testing.
- Strong Python and SQL skills, with experience building maintainable data pipelines, experiment workflows, services, or analytical tools.
- Practical knowledge of information retrieval, embeddings, hybrid search, reranking, classification, structured outputs, tool calling, or common LLM failure modes.
- Strong statistical judgment covering sampling, leakage prevention, uncertainty, calibration, precision and recall, temporal drift, and controlled comparisons.
- Ability to work with semi-structured engineering data from issue tracking, source control, code reviews, CI systems, logs, dashboards, and test systems.
- Clear communication skills and the ability to explain results, limitations, and tradeoffs to technical and business stakeholders.
Se valora
- Knowledge of LTE, 5G NR, Open RAN, or telecom support workflows.
- Experience with enterprise search, RAG evaluation, knowledge graphs, process mining, anomaly detection, or graph-based analysis.
- Familiarity with Jira, Git or Bitbucket, CI/CD telemetry, evaluation frameworks, experiment tracking, or data versioning.
- Experience with open-weight LLMs, commercial model APIs, local inference, proprietary code, customer logs, or access-controlled engineering data.
- Experience handling privacy-sensitive or regulated data in secure enterprise environments.