Data Engineer
Guarda esta oferta y sigue tu búsqueda
Crea una cuenta gratis para guardar empleos, crear alertas y volver a esta oferta desde tu panel.
GROUP BNP PARIBAS
BNP Paribas Group is the top bank in the European Union and a major international banking establishment. It has close to 185,000 employees in 65 countries. In Spain we are more than 5,100 employees within 13 business lines.
IDRAS
Who are we?
We are South Europe Technologies (SET Iberia); the IT, Data and Operations Shared Service Center of BNP Paribas Personal Finance, with delivery centers in Spain and Portugal, providing the best solutions to BNPP PF entities around the world such as Cetelem (specialized, between others, in financial partnership of major retailers, consumer goods companies and car dealerships).
Among Other Services, Our Portfolio Is Composed Of
- Applications Management (Architecture, Project Management, Development, and Quality Assurance)
- IT Risks & Cybersecurity Services
- Platforms Management
- Data Analytics and AI
- Operations
Our offices are in Spain (Madrid) and Portugal (Lisbon, Porto). The company brings together over 200+ employees, with expertise in various technologies (Java, .Net, Python, Tibco, APIGee) and other operational roles (Functional Analyst, Project Manager, Business Analyst, Auto Stock Financing operators). We keep growing
Our consistent track record of services delivery means comfort for our customers and opportunities for our employees.
You will find SE.T to be full of energy and an Inclusive Workplace in which you truly can make a difference.
Would you like to join our international team that delivers end-to-end solutions (applications and operations activities) to businesses of BNP Paribas Personal Finance Group entities around the world?
In a context of maintaining the high level of existing activities while growing the number of international customers, we are looking for our Data Engineer - Developer.
Mission
ABOUT THE JOB
The data engineer is a key role in any of the different data squads in charge of development and maintenance of new big data data platforms and data products. His/her main mission is to develop the different data pipelines which ingest, transform and prepare data from different sources into the different layers of our datahub platform and/or data products.
Responsibilities
- Data Modeling and Pipelines development with Spark and Scala to ingest and transform data from different sources (Kafka topics, APIs, HDFS, structured databases, files...) into HDFS, IBM Cloud Storage (generally in parquet format) or SQL/NOSQL databases
- followign complex business rules.
- Manage big data storage solutions in the platform (HDFS, IBM Cloud Storage, structured and non-structured databases).
- Data Transformation and Quality: implement data transformation and quality control processes to ensure data consistency and accuracy. Utilize programming languages such as Scala and SQL. And libraries like Spark for data transformation and enrichment
- operations.
- CI/CD Pipeline Implementation: set up CI/CD pipelines to automate deployment, unit testing, and development management.
- Infrastructure Migration: migrate the existing Hadoop infrastructure to cloud infrastructure on Kubernetes Engine, Object Storage (IBM Cloud storage), Spark as a service on Scala (to build the data pipelines), and Airflow as a service (to orchestrate and
- schedule the data pipelines).
- Implementation of schemas,queries, and views in SQL/NOSQL databases like Oracle, Postgres or MongoDB.
- Develop and configure scheduling of data pipelines with a combination of shell scripting and AirFlow as a service.
- Validation Testing: conduct unit and validation tests to ensure accuracy and integrity.
- Documentation: write technical documentation (specifications, operational documents) to ensure knowledge capitalization.
- Code Improvement: modify the existing code as per business requirements and continuously improve for better performance and maintainability.
- Configure Dremio Data Virtualization to interface with Parquet or as a way to expose the data in the different data products.
- Performance Optimization and Security: ensure the performance and security of the data infrastructure and follow the best practices of Data engineering.
IT Tools
- Spark on Scala as legacy data pipeline development language.
- Spark as a service on Scala as data pipeline development platform.
- Experience in the design and development of streaming procesess using Spark Streaming, Spark Structure Streaming and Apache Kafka.
- Management of legacy big data storage solutions (HDFS).
- Management of big data storage solutions (IBM Cloud Object Storage and parquet format).
- Implementation of SQL/NO SQL database schemas ,queries and views (MongoDB, Oracle, Postgr