Data Engineer (Mid Level)-Orbit
About Irth Solutions
Irth Solutions is a leading provider of cloud-based SaaS software for damage prevention, asset integrity, stakeholder engagement and land management, helping energy, utility, telecom, and infrastructure companies protect their critical network infrastructure. With nearly three decades of industry experience, Irth serves customers across North America and continues to expand its platform with new data-driven and AI-powered capabilities.
Data Engineer β Insights (AI/ML)
Location: Remote β India
Department: Insights (AI/ML)
Reports to: Data Platform & Analytics Manager
About the Role
Irth is building a modern, multi-cloud, enterprise-grade data estateβa unified Databricks-based data platform that centralizes data across Irth βs products and cloud environments, including AWS, Azure, and GCP .
As a Data Engineer , you will play a hands-on implementation role, working closely with the Senior Data Architect to bring the enterprise data platform vision to life.
You will design and develop data pipelines based on established architectural patterns, implement data quality and governance controls, build Delta Lake and medallion architecture solutions, and help operationalize the new data platform.
This is an excellent opportunity for a mid-level Data Engineer looking to deepen their expertise in Databricks, Apache Spark, cloud data engineering, and modern lakehouse architecture while working in a multi-cloud enterprise environment.
Key Responsibilities
1. Data Pipeline Development β Primary Responsibility
- Build, maintain, and enhance data ingestion pipelines across AWS, Azure, and GCP , following architecture and engineering patterns established by the Senior Data Architect.
- Develop both batch and streaming pipelines using:
- Databricks Workflows
- Apache Spark / PySpark
- SQL
- Delta Live Tables
- Databricks Lakeflow components
- Implement Bronze β Silver β Gold medallion architecture patterns for ingestion, transformation, cleansing, and standardization.
- Implement Change Data Capture (CDC) and Slowly Changing Dimensions ( SCD Type 1 and Type 2 ).
- Handle schema evolution and changing source-system structures.
- Implement data validation, reconciliation, and quality rules as part of pipeline processing.
- Build reusable and maintainable pipeline components following established engineering standards.
2. Platform & Storage Implementation
- Configure and maintain Delta Lake storage structures, tables, schemas, partitions, and optimization routines.
- Apply Delta Lake performance and maintenance practices, including:
- OPTIMIZE
- Z-ORDER
- VACUUM
- Appropriate partitioning and file-management strategies
- Assist with implementation of metadata, cataloging, and lineage standards using Unity Catalog .
- Support integration between cloud storage platforms and Databricks, including:
- Amazon S3 β Databricks
- Azure Storage β Databricks
- Google Cloud Storage β Databricks
- Assist with implementation of scalable storage and processing patterns defined by the Data Architect.
3. Data Governance, Quality & Compliance Enablement
- Implement automated data-quality checks, profiling, validation, and monitoring in accordance with enterprise governance standards.
- Apply data-quality rules at appropriate stages of the Bronze, Silver, and Gold layers.
- Implement RBAC policies, security controls, and data-classification tags defined by the enterprise governance model.
- Support implementation of metadata and lineage mapping across Unity Catalog and Microsoft Purview .
- Help ensure datasets are properly documented, classified, governed, and discoverable.
- Support remediation of data-quality and governance issues identified through monitoring or reviews.
4. Orchestration, Automation & Operational Support
- Build, schedule, monitor, and maintain production workflows using:
- Databricks Workflows
- Delta Live Tables
- Azure Data Factory (ADF)
- Other approved orchestration
Get remote data science jobs like this by email
10 hand-picked jobs, one email a day. No spam, unsubscribe anytime.