Data Engineering Fabric (WFH)

๐Ÿข EXL ยท all 82 jobs
๐Ÿ“ India
๐Ÿ“… Posted Sep 19, 2026 ยท via Internshala
๐Ÿท Remote
Apply on original site โ†—

About the job:

Role Purpose

Build and operate the data pipelines that feed the Entity Hub. This role lands all six in-scope sources into Fabric, implements standardization and transformation logic, and maintains the data quality checks and monitoring that the entity resolution engine depends on. Reliable, observable ingestion is the foundation the entire programme rests on.

Key Responsibilities
- Ingestion development build and maintain pipelines to land the six in-scope sources (Secretary of State, D&B, ARROW, E1, hCue, DocCentral) into the Fabric Bronze/raw layer.
- Mirroring & CDC implement Fabric Mirroring for supported structured sources and establish change-data-capture patterns; implement watermark/incremental load logic where mirroring is unavailable.
- Raw layer management maintain one Delta table per source on an append-only basis, retaining evidence records and full source provenance.
- Standardization & transformation implement name normalization, address parsing and attribute standardization logic in Spark notebooks; support identifier-spine construction.
- Data quality implement data quality checks, validation rules, threshold alerts and exception handling; support reconciliation against source.
- Pipeline operations schedule, monitor and troubleshoot pipeline runs; investigate failures and performance issues; maintain run documentation.
- Performance tuning optimise Spark jobs, Delta file sizes, partitioning and pipeline efficiency to manage Fabric capacity consumption.
- Documentation produce and maintain source-to-target mappings, transformation logic documentation and lineage records.

-

Must-Have Qualifications
- 4+ years hands-on data engineering with strong PySpark and SQL
- Production experience building ingestion pipelines from multiple heterogeneous sources
- Working knowledge of Delta Lake and medallion/lakehouse architecture
- Experience implementing incremental loads and CDC-style processing
- Experience implementing data quality checks and troubleshooting pipeline failures

Nice-to-Have
- Microsoft Fabric hands-on experience (Mirroring, Copy Jobs, Environments)
- Exposure to entity/master data standardization (name and address parsing)
- Familiarity with libraries such as Great Expectations for data quality
- Experience optimising for Fabric capacity/CU consumption

Key Deliverables Owned
- Operational ingestion pipelines for all agreed sources
- Bronze/raw layer with one Delta table per source and CDC retained
- Standardization and parsing transformation logic
- Data quality checks, monitoring and exception handling
- Source-to-target mapping and run documentation

Who can apply:

Only those candidates can apply who:
- have minimum 4 years of experience

Salary:
Competitive salary

Experience:
4 year(s)

Deadline:
2035-01-01 00:00:00

Flights + hotels

This role requires you to be in India. If that means relocating or flying in, it is worth checking fares before you commit to a start date.

Compare flights and hotels โ†’

โ† All remote jobs

Comparing Data Engineer pay and openings โ€” the live median is $128k?All remote Data Engineer jobs โ†’Data Engineer salary data โ†’
Get new remote jobs like this by email
Daily email, only when there's something new. One click to stop.

Get remote jobs like this by email

10 hand-picked jobs, one email a day. No spam, unsubscribe anytime.

Similar for you