Senior DevOps Engineer, CI/CD Platform

Hace 7 horas

Barcelona, Catalonia, España Jobrapido Jornada completa
Overview As a Senior Platform Engineer, you own the CI/CD Kubernetes platform end to end, shaping capacity, reliability, upgrades, security, and cost. You work with a cross-functional engineering team to keep a self-hosted runner fleet fast and scalable for hundreds of developers. You’ll drive performance improvements, manage bare-metal infrastructure, and turn incident learnings into robust runbooks and guardrails. This role blends hands-on infrastructure with product-minded collaboration to make builds faster and more predictable.

Compensaciones / BeneficiosAlan private health insuranceWellhub fitness facilitiesCobee expense savingsLanguage classesOffice breakfast and organic fruitPet-friendly office ResponsabilidadesRightsize runner tiers using weekly CPU/memory data and update deployment via pull requestsInvestigate and fix flaky CI jobs by tracing issues to container runtime, kernel, or clock driftAdd and onboard machines to the runner fleet with provisioning and networkingReduce queue wait times by identifying blockers across pipeline stagesUpgrade clusters or controllers without disrupting usersCollaborate with product engineers to optimize workflows that are slow for a reasonOwn incident response for the platform, run blameless postmortems, and implement alerts and runbooksOperate bare metal infrastructure with focus on Linux, networking, and troubleshooting Requisitos principales5+ years of production infrastructure with Kubernetes at the coreDeep knowledge of Kubernetes scheduling, resources, evictions, node pressure, DaemonSets, controllers and operatorsStrong Linux and container fundamentals (containerd or Docker, cgroups, storage, networking)Terraform and infrastructure-as-code, with modular design and state hygieneCI/CD at platform level; GitHub Actions experience including self-hosted runnersScripting in Bash and at least one of Python or GoObservability with metrics, logs, traces; OpenTelemetry and Prometheus toolingOperational maturity: on-call, incident command, postmortemsCapacity and cost awareness; ability to size a fleet and explain billingClear written English and documentation skillsCollaborative mindsetProblem-solving and ownershipCommunication and ability to translate tech details for engineersKubernetes architecture and operations (scheduling, limits, evictions, DaemonSets, operators)Bare-metal provisioning and troubleshootingGitHub Actions runner platform and runner controller concepts