Senior DevOps Engineer, CI/CD Platform
Hace 7 horas
Barcelona, Catalonia, España
Jobrapido
Jornada completa
Gratis con email o Google
Guarda esta oferta y sigue tu búsqueda
Crea una cuenta gratis para guardar empleos, crear alertas y volver a esta oferta desde tu panel.
Gratis con email o Google
Overview
As a Senior Platform Engineer, you own the CI/CD Kubernetes platform end to end, shaping capacity, reliability, upgrades, security, and cost. You work with a cross-functional engineering team to keep a self-hosted runner fleet fast and scalable for hundreds of developers. You’ll drive performance improvements, manage bare-metal infrastructure, and turn incident learnings into robust runbooks and guardrails. This role blends hands-on infrastructure with product-minded collaboration to make builds faster and more predictable.
Compensaciones / BeneficiosAlan private health insuranceWellhub fitness facilitiesCobee expense savingsLanguage classesOffice breakfast and organic fruitPet-friendly office ResponsabilidadesRightsize runner tiers using weekly CPU/memory data and update deployment via pull requestsInvestigate and fix flaky CI jobs by tracing issues to container runtime, kernel, or clock driftAdd and onboard machines to the runner fleet with provisioning and networkingReduce queue wait times by identifying blockers across pipeline stagesUpgrade clusters or controllers without disrupting usersCollaborate with product engineers to optimize workflows that are slow for a reasonOwn incident response for the platform, run blameless postmortems, and implement alerts and runbooksOperate bare metal infrastructure with focus on Linux, networking, and troubleshooting Requisitos principales5+ years of production infrastructure with Kubernetes at the coreDeep knowledge of Kubernetes scheduling, resources, evictions, node pressure, DaemonSets, controllers and operatorsStrong Linux and container fundamentals (containerd or Docker, cgroups, storage, networking)Terraform and infrastructure-as-code, with modular design and state hygieneCI/CD at platform level; GitHub Actions experience including self-hosted runnersScripting in Bash and at least one of Python or GoObservability with metrics, logs, traces; OpenTelemetry and Prometheus toolingOperational maturity: on-call, incident command, postmortemsCapacity and cost awareness; ability to size a fleet and explain billingClear written English and documentation skillsCollaborative mindsetProblem-solving and ownershipCommunication and ability to translate tech details for engineersKubernetes architecture and operations (scheduling, limits, evictions, DaemonSets, operators)Bare-metal provisioning and troubleshootingGitHub Actions runner platform and runner controller concepts
Compensaciones / BeneficiosAlan private health insuranceWellhub fitness facilitiesCobee expense savingsLanguage classesOffice breakfast and organic fruitPet-friendly office ResponsabilidadesRightsize runner tiers using weekly CPU/memory data and update deployment via pull requestsInvestigate and fix flaky CI jobs by tracing issues to container runtime, kernel, or clock driftAdd and onboard machines to the runner fleet with provisioning and networkingReduce queue wait times by identifying blockers across pipeline stagesUpgrade clusters or controllers without disrupting usersCollaborate with product engineers to optimize workflows that are slow for a reasonOwn incident response for the platform, run blameless postmortems, and implement alerts and runbooksOperate bare metal infrastructure with focus on Linux, networking, and troubleshooting Requisitos principales5+ years of production infrastructure with Kubernetes at the coreDeep knowledge of Kubernetes scheduling, resources, evictions, node pressure, DaemonSets, controllers and operatorsStrong Linux and container fundamentals (containerd or Docker, cgroups, storage, networking)Terraform and infrastructure-as-code, with modular design and state hygieneCI/CD at platform level; GitHub Actions experience including self-hosted runnersScripting in Bash and at least one of Python or GoObservability with metrics, logs, traces; OpenTelemetry and Prometheus toolingOperational maturity: on-call, incident command, postmortemsCapacity and cost awareness; ability to size a fleet and explain billingClear written English and documentation skillsCollaborative mindsetProblem-solving and ownershipCommunication and ability to translate tech details for engineersKubernetes architecture and operations (scheduling, limits, evictions, DaemonSets, operators)Bare-metal provisioning and troubleshootingGitHub Actions runner platform and runner controller concepts