Data Engineer
Data Engineer – Remote
Position Type: Full-Time, Remote
Working Hours: U.S. Client Business Hours (flexibility for pipeline monitoring, deployments, and data refresh cycles)
About the Role
At Pavago , one of our clients is hiring a Data Engineer to design, build, and maintain scalable data infrastructure that powers analytics, reporting, and business decision-making.
You’ll build reliable data pipelines, optimize data warehouses, ensure data quality, and collaborate with engineering, analytics, and business teams to deliver trusted, high-performance data solutions.
What You’ll Do
Data Pipelines & Integration
- Build and maintain ETL/ELT pipelines using Python, SQL, or Scala.
- Orchestrate workflows with Airflow, Prefect, Dagster, or similar tools.
- Integrate data from APIs, databases, SaaS platforms, files, and streaming sources.
- Develop scalable data ingestion workflows.
Data Warehousing & Modeling
- Manage cloud data warehouses such as Snowflake, BigQuery, or Redshift.
- Design scalable data models and schemas.
- Optimize warehouse performance through partitioning, clustering, and indexing.
- Build analytics-ready datasets for reporting and BI.
Data Quality & Governance
- Implement validation, monitoring, and anomaly detection.
- Maintain documentation, lineage, and data governance standards.
- Ensure reliable, audit-ready data processes.
- Monitor pipeline health and resolve failures proactively.
Streaming & Real-Time Processing
- Build and maintain real-time data pipelines.
- Support event-driven architectures and streaming platforms.
- Optimize performance and reliability of streaming workflows.
Collaboration & Analytics
- Partner with analysts, data scientists, and business teams.
- Support reporting initiatives in Power BI, Tableau, Looker, or similar platforms.
- Translate business requirements into scalable data solutions.
- Document pipelines, workflows, and data models.
Infrastructure & Automation
- Deploy data services using Docker and Kubernetes.
- Support CI/CD pipelines and cloud infrastructure.
- Improve system scalability, reliability, and cost efficiency.
Required Experience & Skills
- 3+ years of experience in Data Engineering, Data Infrastructure, or Back-End Engineering.
- Strong Python and SQL skills.
- Experience with Snowflake, BigQuery, Redshift, or similar cloud data warehouses.
- Hands-on experience with Airflow, Prefect, or similar orchestration tools.
- Strong understanding of ETL/ELT pipelines and data modeling.
- Experience with AWS, Azure, or Google Cloud.
Nice to Have
- Experience with dbt.
- Kafka, Kinesis, Pub/Sub, or other streaming platforms.
- AWS Glue, GCP Dataflow, or Azure Data Factory.
- Docker, Kubernetes, Terraform, or CI/CD pipelines.
- Experience in healthcare, fintech, SaaS, or other regulated industries.
- Experience optimizing warehouse performance and cloud costs.
What a Typical Day Looks Like
- Monitor pipeline health and troubleshoot failures.
- Build and maintain data ingestion pipelines.
- Optimize SQL queries and warehouse performance.
- Deliver reliable datasets for analytics and reporting.
- Implement monitoring and data quality checks.
- Document pipelines and data models.
In short: You’ll build and maintain reliable data infrastructure that enables accurate reporting, analytics, and business decisions.
Key Metrics for Success
- 99%+ pipeline uptime.
- Data freshness maintained within SLA targets.
- High data quality with minimal downstream issues.
- Improved warehouse performance and cost optimization.
- Reliable delivery of scalable datasets.
- Strong stakeholder satisfaction.
Interview Process
- Application Review
- Spark Hire Intro Video (3–5 minutes)
- Technical Assessment (ETL Pipeline or SQL Exercise)
- Client Interview
- Offer & Onboarding
What Happens After You Apply
After submitting your application, you’ll receive an email invitation from Spark Hire to record