Data Engineer
Job Summary
We are looking for an experienced Data Engineer to design, build, and maintain a scalable, metadata-driven data ingestion and transformation framework on Google Cloud Platform (GCP) .
The role will be responsible for developing configuration-driven ingestion pipelines, onboarding multiple source systems, and implementing Raw, Bronze, Silver, and Gold data layers . The Data Engineer will use technologies including BigQuery, Cloud Composer/Apache Airflow, Python, Dataform, Datastream, Pub/Sub, Dataflow, Cloud Run, and Cloud Storage to deliver reliable and reusable data solutions.
The ideal candidate will have strong experience in metadata-driven pipeline development, CDC, advanced SQL, Python, data quality, reconciliation, automated testing, and CI/CD . The engineer will work closely with Data Governance and other platform stakeholders to ensure that data pipelines are scalable, governed, auditable, testable, and production-ready.
Responsibilities
-
Design and build a metadata-driven and configuration-driven ingestion framework using reusable pipeline templates.
-
Design, develop, and maintain the metadata configuration model covering source systems, source objects, load configurations, execution parameters, run logs, audit information, and reprocessing.
-
Build dynamic Cloud Composer / Apache Airflow DAGs that generate and execute ingestion workflows based on metadata configurations.
-
Develop reusable ingestion patterns and pipeline generators instead of building separate pipelines for individual source tables.
-
Onboard multiple source systems, including relational databases, SAP, REST APIs, files, streaming sources, and other enterprise platforms .
-
Implement data ingestion into Raw and Bronze layers following defined lakehouse architecture and engineering standards.
-
Implement source-to-target reconciliation to ensure completeness and accuracy of ingested data.
-
Implement robust operational controls, including error handling, retries, quarantine processing, alerting, monitoring, and audit logging .
-
Develop replay-by-batch and reprocessing capabilities to support failed loads, historical reloads, and controlled data recovery.
-
Support batch, incremental, streaming, and Change Data Capture (CDC) ingestion patterns.
-
Implement Bronze-layer processing, including data typing, cleansing, schema validation, and schema enforcement .
-
Implement data de-duplication using appropriate source keys, primary keys, or business keys.
-
Implement soft-delete representation and appropriate handling of deleted source records.
-
Implement CDC change-history materialization and maintain historical changes where required.
-
Develop appropriate BigQuery partitioning, clustering, MERGE, and incremental processing strategies .
-
Perform data compaction and other performance optimization activities where required.
-
Design and develop Silver and Gold transformation models using Dataform .
-
Develop reusable transformation components, macros, dependencies, and incremental transformation strategies.
-
Implement assertions and automated tests for transformation models and ensure no untested transformation reaches production .
-
Implement in-pipeline data quality checks, validation rules, and automated promotion gates .
-
Work closely with the Data Governance Consultant to incorporate data quality, governance, lineage, audit, and control requirements into the engineering framework.
-
Prevent data that fails critical quality requirements from being promoted to downstream layers.
-
Implement geospatial data ingestion , including GeoJSON processing and conversion to BigQuery GEOGRAPHY.
-
Ensure appropriate preservation and handling of Spatial Reference System (SRS) information for geospatial datasets.
-
Implement ingestion and management of semi-structured and unstructured data using Google Cloud Storage and BigQuery object tables .
-
Follow Git-based development practices , includi
This role requires you to be in Pakistan. If that means relocating or flying in, it is worth checking fares before you commit to a start date.
Compare flights and hotels βGet remote data science jobs like this by email
10 hand-picked jobs, one email a day. No spam, unsubscribe anytime.
Similar for you
Get 10 hand-picked remote jobs like this one in your inbox every morning. One email a day, matched to what you browse. No spam, one-click unsubscribe.
No thanks β continue to the application β