AWS Lakehouse Data Engineer

๐Ÿข Guidehouse ยท all Guidehouse jobs
๐Ÿ“ United States
๐Ÿ’ฐ USD 113,000 - 188,000 / annual
๐Ÿ“… Posted 2026-09-03 ยท via Himalayas
๐Ÿท Data-Engineering,AWS-Data-Engineering,Lakehouse-Engineering,ETL-ELT-Pipeline-Development,Cloud-Data-Platform-Engineering,AWS-Data-Engineer,Data-Lakehouse-Engineering,AWS-ETL-Data-Engineer,Data-Lakehouse-Specialist,AWS-Cloud-Data-Engineering,Cloud-Data-Engineer,Data-Engineer
Apply on original site โ†—

Job Family:
Software Development & Support Travel Required:
None Clearance Required:
Ability to Obtain Public Trust AWS Lakehouse Data Engineer

We are seeking an AWS Lakehouse Data Engineer to design, implement, and operate the cloud-native data platform that powers AI/ML, analytics, reporting, and data visualization. You will build a modern lakehouse on Amazon S3 using AWS-native services and open table formats, providing Databricks-like capabilities while maintaining portability, strong governance, cost efficiency, and operational control. You will also develop scalable batch and streaming ingestion, Python and PySpark ETL/ELT pipelines, metadata and governance services, and automated cloud provisioning and CI/CD across environments.

This role is ideal for an engineer who enjoys platform building, automation, performance optimization, and enabling advanced analytics through trusted, secure, and well-governed data.

What You Will Do

Build and Operate Data Pipelines (Batch and Streaming)
- Design and implement batch and streaming ingestion from APIs, relational databases, file drops, event streams, and external partners.

- Implement, test, and optimize ETL/ELT pipelines using Python and PySpark to produce curated, analytics-ready datasets for reporting, visualization, and machine learning.

- Implement incremental processing, change data capture (CDC), data contracts, schema validation, and reusable transformation frameworks.

- Improve pipeline reliability through automated testing, orchestration, monitoring, retry handling, and operational runbooks.

Deliver an AWS-Native Lakehouse Data Platform
- Design and implement a Delta Lakehouse-style data platform using AWS-native services to provide Databricks-like capabilities for data engineering, analysis, and data visualization.

- Build and manage a scalable lakehouse on Amazon S3 using Apache Iceberg and open columnar formats such as Apache Parquet.

- Implement SQL-like table reliability for data stored in Amazon S3, including ACID transactions, schema evolution, partition evolution, snapshot isolation, time travel, and rollback capabilities using Apache Iceberg.

- Enable fast, interactive querying of lakehouse data using AWS-native query and compute services such as Amazon Athena, Amazon EMR, AWS Glue, and Amazon Redshift where appropriate.

- Optimize performance and cost through partitioning, compaction, file sizing, statistics, caching, lifecycle policies, and efficient separation of compute and storage.

- Establish standardized development, test, and production environments with consistent configuration and controlled promotion across stages.

Metadata, Governance, Access Control, Lineage, and Quality
- Implement data governance and fine-grained access control using AWS-native services, including AWS Lake Formation, AWS Glue Data Catalog, AWS Identity and Access Management (IAM), AWS Key Management Service (KMS), and related security services.

- Implement a managed metadata repository for dataset cataloging, ownership, business definitions, tagging, classification, and discoverability.

- Enable end-to-end lineage from source through transformation and consumption to support auditability, impact analysis, and regulatory requirements.

- Apply policy-based access, least-privilege permissions, row-, column-, and cell-level controls where required, data classification, retention, encryption, and secure data handling.

- Build operational data quality checks for freshness, completeness, uniqueness, validity, consistency, and anomaly detection, and publish measurable SLAs/SLOs.

AWS Automation, CI/CD, and Operations
- Implement automated AWS provisioning using Infrastructure as Code (IaC) to create consistent environments and secure-by-default baselines.

- Build and enhance CI/CD for data pipelines and lakehouse components, including automated tests, security checks, validation gates, packaging, deployment, promotion, and rollback strategies.

โ† All remote jobs

Similar for you