Lead Data Engineer

Hace 3 semanas

Cataluña, Cataluña, España dsm-firmenich Jornada completa

Job Title: Lead Data Engineer (Cross Domain)

Location: Barcelona, Spain

Join us as a Lead Data Engineer to design and maintain robust data pipelines, drive the best deployment practices, and manage key initiatives in the Procurement domain. You will own the full lifecycle of data products — from source ingestion to Gold layer consumption — using modern DevOps practices across Azure DevOps and GitHub.

You will collaborate with data modelers, data stewards, and business partners to deliver impactful, production-grade data solutions. Elevate your career by mastering cutting-edge tools — dbt, Databricks, PySpark, and CI/CD automation — in a dynamic, innovative, and global environment.

Your key responsibilities

  • Design and implement end-to-end data pipelines for ingestion, transformation, and storage across Bronze (raw), Silver (Raw Vault and Business Vault), and Gold (data marts and semantic layer), ensuring scalable and reliable data processing.
  • Develop and manage data ingestion processes from source systems through ETL and API-based extraction methods where platform tables are unavailable, ensuring seamless integration of enterprise data sources.
  • Build modular, reusable, and maintainable transformation frameworks using dbt, PySpark, and SQL, including the creation of staging models (raw_stage and stage) with hash keys, hash diffs, business keys, and load metadata in alignment with Data Vault 2.0 and modern engineering best practices.
  • Design, build, and maintain CI/CD and DevOps processes in Azure DevOps and GitHub Actions, incorporating automated testing, linting, deployment gates, environment promotion, and standardized branching, pull request, code review, and merge strategies.
  • Implement and manage Infrastructure-as-Code (IaC) and deployment automation for data platform resources, including provisioning environments, deploying dbt jobs and Databricks workflows, configuring clusters, managing dependencies, scheduling jobs, and maintaining environment‐specific configurations.
  • Establish proactive monitoring, observability, and operational excellence practices by implementing alerting, conducting regular pipeline maintenance and upgrades, and troubleshooting issues to ensure platform reliability, performance, and rapid incident resolution.
  • Drive data quality, governance, and compliance standards through automated dbt testing, validation frameworks, FAIR data principles, lineage management, auditability, and data discoverability across all layers of the platform.
  • Provide technical leadership and cross‐functional collaboration by evaluating solution feasibility, leading workload planning, mentoring data engineering teams, defining engineering standards, and partnering with data engineers, modelers, stewards, BI developers, data scientists, and business SMEs to deliver impactful data solutions.

We offer

  • Unique career paths across health, nutrition, and beauty — explore what drives you and get the support to make it happen.
  • A chance to impact millions of consumers every day — sustainability embedded in all we do.
  • A science‐led company with cutting‐edge research and creativity everywhere — from biotech breakthroughs to sustainability game‐changers, you'll work on what's next.
  • Growth that keeps up with you — you join an industry leader that will develop your expertise and leadership.
  • A culture that lifts you up — with collaborative teams, shared wins, and people who cheer each other on.
  • A community where your voice matters — your ideas are essential to our future.

You bring

  • Strong technical expertise in dbt, SQL, Python, Spark (PySpark), Databricks, Git, Azure DevOps, and GitHub, with proven experience designing, building, and maintaining production‐grade data pipelines for 5+ years.
  • Deep CI/CD and automation experience, including the design and implementation of deployment pipelines using Azure DevOps Pipelines and GitHub Actions, with Infrastructure-as-Code knowledge (Terraform, Bicep) considered a strong advantage.
  • Advanced Git and repository management skills, including branching strategies, pull request standards, code review best practices, and governance across Azure DevOps and GitHub environments.
  • Strong cloud data engineering expertise (Azure preferred), with hands‐on experience deploying, managing, and operating Databricks jobs, clusters, and workflows in production environments.
  • Extensive knowledge of Data Vault 2.0 architecture, including Raw Vault (Hubs, Links, Satellites), Business Vault components (bSAT, PIT, Bridge), and Gold‐layer data marts and semantic models.
  • Expertise in data ingestion and integration patterns, including batch ETL, incremental processing, Change Data Capture (CDC), and API‐based ingestion from a wide range of source s