AI Engineer
Guarda esta oferta y sigue tu búsqueda
Crea una cuenta gratis para guardar empleos, crear alertas y volver a esta oferta desde tu panel.
Al continuar, aceptas nuestros Términos & Política de Privacidad.
AI EngineerDepartment: DeliveryEmployment Type: Full TimeLocation: MadridDescriptionAt Sabio Group, we're building the next generation of AI-powered customer experience for some of the world's most demanding enterprise brands — the kind of frontier customers who set the bar for what "great CX" looks like. We pair frontier-class large language models with deep contact centre expertise to deliver agentic experiences that materially improve containment, CSAT and operational efficiency.
We're hiring an AI Engineer to join our customer service bot / AI agent team in Madrid. You'll work end-to-end across the lifecycle — discovery, design, build, evaluate, deploy, operate — engineering robust LLM/NLU solutions, knowledge pipelines and integrations across voice and digital channels, primarily in Spanish.
This is a hands-on engineering role for someone who is curious about where AI is going, comfortable shipping agentic systems into production, and thoughtful about doing so safely and cost-effectively.
Key Responsibilities
Discovery & DesignTranslate CX and business needs into AI solution designs spanning NLU/LLM, RAG, ASR/TTS and back-end integrations.
Define functional and non-functional
requirements:
latency, resiliency, observability, security, cost-per-conversation.
Agent & Model ImplementationDesign and build agentic AI solutions — single-agent, multi-agent, tool-using and orchestrated patterns — using modern frameworks and frontier models.
Build and optimise NLU/LLM components: classification, NER, summarisation, RAG over enterprise knowledge bases.
Engineer prompts, tool schemas, context windows and token budgets for accuracy, latency and cost.
Develop data and knowledge pipelines for ingestion, cleansing, PII redaction and evaluation datasets.
Evaluation & QualityBuild evaluation harnesses (offline eval sets, LLM-as-judge, regression suites, A/B testing) and treat eval as a first-class deliverable, not an afterthought.
Drive continuous improvement via conversation analytics, error triage and failure-mode analysis.
LLMOps & OperationsImplement CI/CD for models and prompts, with feature flags, canary releases and rollback.
Track experiments, lineage, and model/prompt versions.
Operate services in production against SLOs with monitoring, tracing, alerting and incident response.
Responsible AIApply practical guardrails for safety, bias, hallucination, prompt-injection and jailbreak resistance.
Enforce GDPR / LOPDGDD: consent, minimisation, retention, access control, auditability.
CollaborationWork alongside conversation designers, software engineers, data scientists, QA and Ops.
Produce clear design docs, runbooks and stakeholder updates.
Skills Knowledge and ExpertiseRequiredHands-on experience delivering conversational AI — voicebots, chatbots or virtual assistants — including conversational flows, open-ended interactions, NLU and generative AI.
Production experience building agents with generative AI: agentic architectures (single-agent, multi-agent, tool/function-calling, planner–executor patterns), RAG, and orchestration.
Working knowledge of frontier LLM providers and platforms — e.g. Anthropic Claude, OpenAI, Google Gemini, Amazon Bedrock, Microsoft Azure AI — and a practical sense of when to use which.
Strong prompt engineering and context engineering skills across multiple LLM families, including system prompt design, structured output, and grounding strategies.
An evaluation mindset: you measure agent quality with eval sets, traces and metrics — not vibes.
Awareness of AI ethics, bias and safety, with practical experience mitigating them in deployed systems.
Token, latency and cost optimisation — model routing, caching (incl. prompt caching), context compression, retrieval tuning — as a core engineering discipline.
Comfort with agentic development workflows — using AI coding assistants and AI co-work / pair-development models (Claude Code, Copilot, Cursor or equivalent) as part of your day-to-day delivery.
Solid software engineering fundamentals: Python and/or JavaScript/Node, Git, IDEs, Agile delivery.
Customer-facing
skills:
requirements gathering, design, validation and stakeholder communication.
Working proficiency in Spanish and English.
DesirableCloud services, primarily AWS, with containerised deployment via Docker / Kubernetes.
Contact centre integrations (e.g. Genesys, Avaya, Cisco, Amazon Connect) and Meta / WhatsApp Business / digital channel integrations.
Experience with ASR/TTS and voice-specific concerns: barge-in, turn-taking, ASR error recovery, latency budgets.
Speech/NLU recognition models and continuous-improvement loops (Microsoft STT, MS CLU, Google, Amazon).
Observability for LLM systems: tracing, eval pipelines, online quality monitoring.
Familiarity with emerging agent interoperability standards (e.g. MCP, A2A) and human-in-the-loop / escalation patterns.
Microsoft Bot Framework, Genesys Dialog Engine or VXML exposure.
Consulting experience analysing existing bot estates, evaluating KPIs (containment, CSAT, AHT