Data Engineer
Data Engineer (Python, SQL, ETL, Airflow, Snowflake & BigQuery) โ Remote
Position Type: Full-Time, Remote
Working Hours: U.S. Business Hours
About the Role
At Pavago , one of our clients is hiring a highly technical Data Engineer to build, maintain, and optimize scalable data pipelines, cloud data infrastructure, and analytics-ready datasets.
Youโll be responsible for the systems that move, transform, validate, and organize data across the business โ ensuring analysts, data scientists, engineering teams, and leadership have access to accurate, reliable, and timely data .
This is a hands-on engineering role focused on ETL/ELT development, data warehousing, SQL optimization, orchestration, data quality, and cloud infrastructure .
If you enjoy building robust data systems, solving complex pipeline problems, and designing infrastructure that scales, this role is for you.
What Youโll Own
ETL / ELT Pipeline Development
- Build and maintain scalable ETL/ELT pipelines using Python and SQL.
- Ingest and process data from:
- APIs
- SaaS platforms
- Relational databases
- Cloud applications
- Streaming systems
- Develop reliable extraction, transformation, loading, and validation workflows.
- Build reusable connectors and pipeline components.
- Troubleshoot pipeline failures and data inconsistencies.
Workflow Orchestration & Automation
- Build and manage workflows using Apache Airflow, Prefect, Dagster, Luigi, or similar tools.
- Monitor pipeline health, scheduling, dependencies, and failed jobs.
- Implement automated retries, alerts, and failure handling.
- Improve pipeline reliability and reduce manual intervention.
- Maintain dependable data freshness across critical datasets.
Data Warehousing & Modeling
- Design and optimize cloud data warehouses using:
- Snowflake
- BigQuery
- Redshift
- Build analytics-ready data models and warehouse structures.
- Develop star and snowflake schemas where appropriate.
- Optimize SQL queries and warehouse workloads.
- Improve performance through partitioning, clustering, indexing, and efficient data modeling.
- Monitor and optimize warehouse costs.
Data Quality & Governance
- Implement automated data validation and quality checks.
- Build monitoring for anomalies, missing data, and transformation failures.
- Maintain logging, lineage, and auditability across pipelines.
- Use tools such as dbt and Great Expectations.
- Establish consistent naming conventions and transformation standards.
- Support governance and compliance requirements, including GDPR, HIPAA, or industry-specific standards where applicable.
Streaming & Real-Time Data
- Build and maintain streaming or event-driven data pipelines.
- Work with technologies such as:
- Kafka
- Kinesis
- Pub/Sub
- Support real-time ingestion and low-latency analytics use cases.
- Ensure streaming workflows remain reliable and scalable.
Cloud Infrastructure & DevOps
- Containerize data services using Docker.
- Support Kubernetes-based environments where applicable.
- Build and maintain CI/CD workflows using GitHub Actions, Jenkins, GitLab CI, or similar tools.
- Support infrastructure-as-code using Terraform or CloudFormation.
- Improve deployment reliability, scalability, and automation across the data platform.
Cross-Functional Collaboration
- Partner closely with Data Analysts, Data Scientists, BI teams, Product, and Engineering.
- Deliver curated datasets for:
- Dashboards
- Business intelligence
- Analytics
- Machine learning
- Operational reporting
- Support BI platforms including Tableau, Looker, and Power BI.
- Maintain clear documentation for pipelines, schemas, workflows, and data definitions.
Required Experience & Skills
- 3+ years of professional Data Engineering or backend engineering experience.
- Strong proficiency in:
- Python
- SQL
- Hands-on experience with at least one modern cloud data warehouse:
- Snowflake
- BigQuery
- Redshift
- Experience bu