DevOps Lead

Hace 4 días

Barcelona, Cataluña, España Allianz Technology Jornada completa

About The Job

We are looking for a permanent, end-to-end DevOps Lead to own DevOps strategy and governance for Allianz in a Pocket, Allianz's flagship global mobile app, across the program's layers. You will align practices already emerging across these teams into a single, coherent strategy, combining direct oversight of a core DevOps function with cross-team influence, working closely with each team's Principal DevOps Engineer to turn program-wide standards into local practice. In scope: the program's KPI framework and unified dashboards; observability, alerting, capacity planning and infrastructure cost monitoring; CI/CD, GitOps and environment strategy; tooling standardization; end-to-end incident management; disaster recovery as accountable owner; release, change, security and risk governance — including threat modeling, access lifecycle management and DevOps-owned control compliance within Allianz's risk framework (Risk Shield/ARS).

What You Do

  • DevOps Strategy, Governance & Tooling Standardization: Define and own the end-to-end DevOps strategy across teams; standardize tooling choices (CI/CD platforms, IaC, observability stack) and align existing team-level practices into a single program-wide way of working, in partnership with each team's Principal DevOps Engineer.
  • Security & Risk Governance: Own DevOps accountability within Allianz's risk and control framework (Risk Shield/ARS) for DevOps-owned controls — e.g. test-before-go-live, production access lifecycle (provisioning, periodic access review, revocation), DR testing evidence; coordinate threat modeling exercises for shared platform components with security and engineering teams.
  • CI/CD Pipeline & Environment Governance: Define program-wide CI/CD standards — pipeline tiering (PR gate/nightly/release candidate), quality gates, deployment strategies — and own the program's environment strategy (dev/test/staging/prod): provisioning standards, parity and lifecycle management across teams.
  • Observability, Alerting, Capacity & Cost Monitoring: Own program-level observability strategy — monitoring, alert thresholds, escalation paths — plus capacity planning (anticipating scaling needs ahead of peaks and OE rollouts) and infrastructure cost monitoring across teams, surfacing capacity, cost and scaling signals to the Backend Lead and engineering teams, who own remediation actions.
  • KPI Framework, Dashboards & SLA Reporting: Define and own the program's DevOps KPIs (deployment frequency, MTTR, change failure rate, availability, SLO attainment); align dashboards across teams into a consistent end-to-end view; define and track SLAs and report platform health to program leadership.
  • Incident Management (E2E): Own end-to-end incident management for production issues — triage, severity classification, escalation, cross-team coordination, and post-mortem process — with the direct support of each team's Principal DevOps Engineer during active incidents.
  • Disaster Recovery — Accountable Owner: Own disaster recovery end-to-end: define scope, RTO/RPO targets and test cadence, and coordinate the teams to ensure DR mechanisms are in place and DR tests are executed and passed across layers and, Release & Change Governance defining release and change management standards (deployment windows, rollback strategy, pre/post-deployment checks) applied consistently across all program layers.
  • AI-Assisted DevOps Practices: Champion the adoption of AI-assisted tooling (e.g. GitHub Copilot, Claude Code) across DevOps workflows to improve automation, troubleshooting and incident response efficiency.

What You Bring

  • DevOps Strategy & Governance Leadership 8+ years in DevOps/SRE/Platform Engineering with at least 3 years leading DevOps strategy and governance across multiple teams. Proven ability to align DevOps practices, KPIs, and dashboards across engineering teams with different technical stacks and maturity levels.
  • Incident Management & Disaster Recovery at Scale Track record owning end-to-end incident management for production systems (triage, escalation, cross-team coordination, post-mortems) in regulated or high-availability contexts. Experience defining and executing disaster recovery strategy including RTO/RPO definition and DR testing across distributed environments.
  • Cross-Team Leadership & Stakeholder Management Demonstrated ability to lead DevOps outcomes across multiple teams and stakeholders without full formal authority. Strong track record coordinating with distributed, multi-country engineering teams across different time zones.
  • Observability, Capacity Planning & Cost Management Hands-on experience with observability tooling (dashboards, alerting, SLOs), capacity planning ahead of scaling events, and infrastructure cost monitoring across cloud environments. Familiarity with FinOps practices and cloud cost optimizati