Data Engineer
Guarda esta oferta y sigue tu búsqueda
Crea una cuenta gratis para guardar empleos, crear alertas y volver a esta oferta desde tu panel.
Al continuar, aceptas nuestros Términos & Política de Privacidad.
Job Description
Kleinanzeigen reaches 36M+ monthly users and is Germany's leading online classifieds market. We are a sustainability-focused platform where you won't just write code;
you will power the services that bridge our high-traffic backend with seamless user experiences.
As a Data Engineer (Analytics & Products) at Kleinanzeigen, you start where the data need begins — not where the code does. You will work directly with product managers, analysts, and domain stakeholders to understand what decisions the data needs to support, how it will be consumed, and what already exists in the platform before writing a single line. From there, you design and build data assets that are homogeneous with the existing stack, cost-effective at scale, maintainable by the next engineer, and self-discoverable, so consumers can find, understand, and trust what you built without asking you.
You will own the full lifecycle: ingestion through Airflow DAGs, transformation across dbt stage, core, and report layers on Databricks, and the quality, documentation, and observability that make those assets production-grade. You stay current with Airflow and design DAG topology deliberately, knowing when to use dynamic task mapping, data-aware scheduling, the TaskFlow API, or a Cosmos dbt task group with Watcher mode for performance-critical pipelines.
You understand Spark well enough to know why a dbt incremental model running in Databricks produces a full table scan instead of a partition filter push-down, how a poorly configured merge operation compounds into file fragmentation over time, and what to do about it, whether that means adjusting the incremental strategy, adding liquid clustering, running OPTIMIZE, or rethinking the model grain entirely.
You write Python and shell scripts as naturally as SQL and follow engineering core principles: modularity, idempotency, testability, not because they are rules, but because they make your work last.
You operate in an AI-First engineering model: AI handles execution;
you own intent, precision, and correctness. You use it to automate, to accelerate best practices, and to raise the quality bar across the team, not to ship faster with less judgment. You will collaborate closely with analysts, product managers, data platform engineers, and stakeholders across the company to ensure that the datasets we create and maintain in our data platform are reliable, trustworthy, and built to serve the decisions that matter.
What You Will Do Data Modelling & dbt Development
- Implement dbt models across the medallion architecture applying the right materialisation strategy for each layer and use case — incremental, full refresh or snapshots — with consistent naming conventions, YAML documentation, metadata tagging, unit tests to validate critical business logic, and reusable macros for common transformation and data replication patterns
- Build and refactor models for different business areas of the company, including for example marketing performance, product metrics, C2C transactions, monetisation, vibrancy, and trust & safety
- Author reusable macros and apply consistent naming conventions, YAML documentation, and metadata tagging (billed, retention, gdpr)
- Design data models that are homogeneous with the existing stack, built at the right grain for the use case, cost-effective to run, and self-discoverable without needing the author to explain them
- Perform cost-aware modelling: clustering strategies, warehouse sizing, incremental scan reduction, and pre-aggregation layers
Data Ingestion & Pipeline Engineering
- Build and maintain Airflow DAGs using Python operators, designing DAG topology deliberately by choosing execution patterns, dependency structures, and sensor logic that match the operational requirements of each pipeline
- Design sensor logic for pipeline dependencies, including intraday vs daily completeness checks and DST-aware temporal handling
- Operate Cosmos dbt task groups within Airflow, including DAG splitting, warehouse selection, and Cosmos version upgrades
- Integrate new data sources end-to-end by building reliable, fault-tolerant connections across heterogeneous endpoint types including REST APIs, event streams, database connectors and file-based sources, with error handling, retry logic, and security patterns that make each integration production-safe from day one, unit testing every operator and transformation component where possible, and defining SLAs and SLOs that set clear expectations on data freshness, completeness, and availability for downstream consumers
- Work with data in the right format for each layer: Avro for event-driven ingestion schemas, Parquet for efficient columnar storage, and Delta for ACID-compliant lakehouse tables with time travel and schema evolution
Data Quality & Reliability
- Write dbt tests (