Data Engineer (Mid Level)

🏢 Irth · all 12 jobs
📍 United States
📅 Posted Sep 13, 2026 · via Himalayas
🏷 Data Engineering, ETL Development, Cloud Data Platform Engineering, Pipeline Engineering, Data Platform Engineer, Mid Level Data Engineer +2 more
Apply on original site ↗

Data Engineer – Insights (AI/ML)

Location: Remote (US)
Department: Insights (AI/ML)
Reports to: Engineering Manager
About the Role

Irth is building a new AI-driven threat and risk management platform for pipeline asset integrity. The platform brings together three capabilities that have historically been separate at Irth :

- A governed, cross-product data platform built on Databricks and Azure

- An AI-powered ingestion layer that normalizes, repairs, and enriches customer data without services-heavy onboarding

- A reusable analytical layer that runs industry-standard, Irth -developed, and customer-built risk models against the data

As a Data Engineer , you will build the pipelines that make this platform real. Pipeline operators hold their integrity data across inline inspection reports, GIS systems, maintenance records, spreadsheets, scanned documents, and enterprise systems of record. Getting that data ingested, cleaned, aligned, and made model-ready is one of the biggest obstacles to adoption in this market—and it is the problem this role exists to solve.

You will implement ingestion and transformation pipelines based on patterns established by the Data Architect, build the AI-assisted ingestion layer in partnership with data scientists, and help operationalize both. This role is well suited to a mid-level engineer who wants to deepen their expertise in Databricks, Spark, and modern lakehouse engineering while helping build a platform from the ground up.
Key Responsibilities
1. Pipeline Development — Primary Responsibility

- Build and maintain ingestion pipelines for structured and semi-structured sources, including GIS, inline inspection data, SCADA, maintenance systems, and enterprise systems of record.

- Implement batch and streaming ingestion using Databricks Workflows, Spark, PySpark, SQL, and declarative pipeline tooling.

- Apply medallion architecture patterns (Bronze, Silver, Gold) for transformation, standardization, and enrichment.

- Implement change data capture (CDC), slowly changing dimensions (SCD), schema evolution, and data-validation rules.

- Normalize third-party and public data feeds, including weather history, soil characteristics, satellite-derived data, and one-call ticket data, into the shared data model.

2. AI-Assisted Ingestion Layer

- Build pipelines that automate normalization of units, schemas, and semantics across inconsistent customer data.

- Implement automated data-quality repair workflows, including gap filling, error correction, and reconciliation, with clear provenance for every synthesized value.

- Work with data scientists to productionize document-extraction pipelines that parse reports, spreadsheets, and field records into the target schema.

- Build human-in-the-loop review and exception workflows so low-confidence extractions are surfaced rather than propagated silently.

3. Platform & Storage Implementation

- Configure and manage Delta Lake tables, partitioning strategies, and optimization routines.

- Implement metadata, lineage, and cataloging standards using Unity Catalog.

- Build and maintain connectors to customer systems of record with configurable refresh cadences.

- Support geospatial data processing, including spatial joins and alignment of results to pipeline centerline geometry.

4. Governance, Quality & Compliance Enablement

- Implement data-quality tests, profiling, and drift monitoring based on standards established by the Data Architect.

- Apply access-control policies, security rules, and classification tags defined by the governance model.

- Implement lineage capture sufficient to support regulatory traceability from ingestion through model output.

5. Orchestration, Automation & Operational Support

- Build, schedule, and monitor data workflows, and own alerting and failure handling for the pipelines you develop.

- Contribute to CI/CD for pipeline code, including version control, automated testing, and environment promotion.

- Tro

← All remote jobs

Want more like this? Browse every live remote data science role.All remote data science jobs →

Get remote data science jobs like this by email

One weekly digest. No spam, unsubscribe anytime.

Similar for you