Senior Site Reliability Engineer
Hace 5 días
La Mancha comarca, Provincia de Ciudad Real; Castilla-La Mancha, España
Stellar Cyber
Jornada completa
Gratis con email o Google
Guarda esta oferta y sigue tu búsqueda
Crea una cuenta gratis para guardar empleos, crear alertas y volver a esta oferta desde tu panel.
Gratis con email o Google
~ We are seeking a highly skilled Senior Site Reliability Engineer (SRE) to join our team and drive reliability, scalability, and efficiency across our production systems.
~ The ideal candidate will have deep expertise in cloud infrastructure, Kubernetes administration, observability, and incident management, with a proven track record of building and maintaining highly available and resilient platforms.
~ As a senior member of the SRE team, you will not only operate complex distributed systems but also influence architecture, tooling, and best practices to ensure operational excellence
~ Administer and maintain container orchestration platforms and containerized workloads
~ Monitor and troubleshoot production systems, participating in on-call rotations to ensure reliability
~ Drive observability improvements by enhancing monitoring, logging, and alerting capabilities across systems and data platforms
~ Administer and optimize cloud-based environments across multiple providers
~ Manage and support distributed data platforms and real-time processing systems
~ Develop and maintain continuous integration and deployment pipelines for efficient and reliable deployments
~ Own and implement Infrastructure as Code (IaC) practices to ensure consistency and scalability
~ Automate and orchestrate infrastructure using programming and scripting languages
~ Perform system administration and networking tasks to support internal and external environments
~ Collaborate effectively with engineers and stakeholders across different time zones Benefits
~ Pre-IPO Stock Options
~ Medical, Dental & Vision care
~401(k)
~ Employee Assistance Program
~ Employee Discount Program
~ Life Insurance
~ Paid time off
~ Referral Program
~ Rewards and Recognition Program
Expertise in operating data platforms (Elasticsearch, MongoDB, Spark, Kafka, Redis)Deep understanding of Infrastructure as Code (Terraform, Helm)Proficiency with public cloud services (AWS, Azure, GCP, or OCI)Hands-on experience with observability tools: Prometheus, Grafana, Loki, and AlertmanagerStrong programming and automation skills in Python and BashCertifications in AWS, GCP, Observability, Linux or Kubernetes are a plus5+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering rolesBachelor's degree in Computer Science, Engineering, or a related technical fieldKnowledge in chat-based operations interfaces and/or auto-remediation controllers using AI agentic frameworkProven success leading large-scale production systems in cloud environments (AWS, GCP, Azure, or OCI)Understanding of AI agents for Auto-triaging alerts, correlate signals and suggest/root-cause hypothesesStrong experience with production on-call operations and incident managementExperience with CI/CD pipelines (GitHub Actions, Bitbucket, ArgoCD)Excellent problem-solving, communication, and leadership abilitiesAdvanced proficiency in Kubernetes administration and troubleshootingStrong technical background in distributed systems, databases, networking, and Linux administrationDemonstrated leadership in driving incident response, on-call best practices, and reliability-focused culture