Senior Infrastructure

Hace 1 día

Barcelona, Catalonia, España PSD Group Jornada completa

Senior Infrastructure & Storage Engineer

Descubra exactamente qué habilidades, experiencia y cualificaciones necesitará para tener éxito en este puesto antes de enviar su solicitud a continuación.
Summary
Location: Barcelona (Hybrid)
Rate: Negotiable
Duration: 6 Months
About the Client
Our client is a global leader in aviation technology, delivering innovative IT and communications solutions that connect airlines, airports, aircraft and governments worldwide. Operating in over 200 countries and territories, they play a critical role in enabling the safe, secure and efficient movement of millions of passengers every day.
About the Role
The Senior Infrastructure & Storage Engineer owns the implementation, integration, testing, and engineering lifecycle of on-premises compute, Linux, distributed storage, and related data-centre infrastructure.
The role translates approved architecture into reliable, resilient, supportable platforms and provides deep engineering support for complex infrastructure incidents.
This is a senior, hands-on role with particular emphasis on bare-metal infrastructure, Ceph, MinIO, Linux, Kubernetes storage integration, capacity, backup and recovery, and operational knowledge transfer. The engineer works closely with Platform Architecture, Operations, Kubernetes & Platform Engineering, Network Engineering, Security, vendors, and delivery teams
Key Responsibilities:
On-Premises Infrastructure Engineering
Build, configure, troubleshoot, and lifecycle-manage physical and virtual servers across heterogeneous data-centre environments.
Engineer and support CentOS, RHEL, or equivalent Linux platforms, including operatingsystem, service, package, certificate, disk, memory, process, and performance issues.
Execute data-centre refreshes, hardware deployments, firmware and operating-system changes, and replacement activities in accordance with approved designs.
Validate compute, network, storage, power, capacity, resilience, monitoring, and support dependencies before production acceptance.
Support selected Windows infrastructure where required.
Distributed Storage and Data Protection
Administer and troubleshoot Ceph, including cluster health, OSDs, monitors, placement groups, pools, capacity, performance, replication, and failure domains.
Engineer and support distributed MinIO deployments, object-storage availability, capacity, healing, replication, and recovery.
Diagnose storage latency, degraded redundancy, disk and node failures, data-path issues, and capacity risks using an evidence-based approach.
Plan, document, and test backup, restoration, disaster recovery, and data-protection procedures; distinguish platform redundancy from backup.
Coordinate specialist and vendor escalation for high-risk or complex storage problems
Kubernetes Infrastructure Integration
Support bare-metal Kubernetes nodes and their operating-system, hardware, network, containerruntime, and storage dependencies.
Engineer and troubleshoot Kubernetes persistent storage, CSI drivers, StorageClasses, persistent volumes, mounts, and Ceph-backed workloads.
Collaborate with the Kubernetes & Platform Engineer during cluster upgrades, node maintenance, capacity changes, recovery testing, and stateful workload incidents.
Understand core Kubernetes concepts sufficiently to diagnose whether failures originate in the workload, node, network, CSI, storage, or underlying infrastructure.
Automation, Monitoring, and Capacity
Automate repeatable infrastructure provisioning and configuration using Terraform, Ansible, Bash, Python, or equivalent tools.
Use Git-based workflows so infrastructure changes are reviewable, auditable, and reproducible.
Implement Prometheus metrics, dashboards, and actionable alerts for hosts, hardware, Ceph, MinIO, storage paths, capacity, and backup health.
Track utilization, saturation, growth, hardware health, and redundancy; implement shortand mediumterm capacity actions.
Integrate infrastructure logs and telemetry with platforms such as New Relic and Elasticsearch where appropriate.
Testing, Handover, and Engineering Support
Plan and execute functional, integration, resilience, failover, capacity, backup, restoration, upgrade, and rollback testing.
Define acceptance criteria, evidence, abort conditions, maintenance procedures, and safe recovery paths.
Provide Level 3 support, lead root-cause analysis, and implement permanent corrective actions for complex infrastructure incidents.
Produce practical SOPs, runbooks, diagrams, asset and dependency records, maintenance procedures, and escalation paths.
Conduct hands-on knowledge transfer and ensure routine operations can be performed safely without dependency on one engineer.
What we are looking for
Required Skills
Strong recent, hands-on experience engineering and troubleshooting produ