Senior Data Engineer

🏢 ParetoHealth · all ParetoHealth jobs
📍 United States
📅 Posted 2026-08-21 · via Himalayas
🏷 Data-Engineer,Data-Engineering,AWS-Data-Engineering,MLOps-Engineering,Data-Pipeline-Engineering,Senior-Data-Engineering,Senior-Data-Engineer-Jobs,Senior-Data-Engineer-Positions
Apply on original site ↗

About ParetoHealth

ParetoHealth is redefining the way employers fund healthcare. As the largest and fastest-growing benefits captive in the United States, we help thousands of small and midsize employers take control of healthcare costs through a smarter, more sustainable model.

Our mission is simple: give small and midsize employers the scale and protection they need to eliminate volatility and lower healthcare costs.

By combining data-driven insights, innovative risk management, and the collective purchasing power of our community, we enable employers to reduce volatility, improve long-term outcomes, and reinvest savings into their businesses and their people.

Headquartered in Philadelphia, ParetoHealth is growing rapidly and transforming one of the country's largest industries. Our success is fueled by talented people who are united by our four core values: Fire in the Belly, For the Greater Good, See the Field, and Get It Done Right. These values shape how we innovate, collaborate, make decisions, and deliver exceptional results for our clients and one another.

If you're energized by solving complex challenges, thrive in a high-growth environment, and want to help reshape the future of healthcare, we'd love to meet you.

Please note that ParetoHealth does not provide employment visa sponsorship for this position. Candidates must be authorized to work in the United States without sponsorship both now or in the future.

Position Summary:

The Senior Data Engineer will design, build, and maintain serverless data pipelines and data models on AWS that make high-quality, analytics-ready data available to our AI and Analytics teams. This role will own ingestions, transformation, and storage patterns using services such as Lambda, Glue, Athena, and S3, ensuring data is reliable, well-documented, and aligned with business and model-training needs. The role will also partner with Data Science to support MLOps, including reproducible machine-learning training and scoring pipeline, automated feature generation, deployment, and monitoring. It will power ML model workflows using sensitive healthcare data and support model training and scoring, production monitoring, and business feedback loops.
Key Responsibilities:

- Design, implement, and maintain scalable, fully serverless data pipelines on AWS using Lambda, Glue, Athena, Step Functions, and S3 to support reporting, analytics, and AI use cases.

- Build and evolve data models and schemas that enable performant querying and downstream consumption by AI, Analytics and engineering teams. Build versioned data models, feature stores, with semantics, reproducible backfills, and safeguards against leakage or inconsistent definitions.

- Develop ETL/ELT workflows to ingest, cleanse, transform, and load data from internal applications, third‑party sources, and event streams into our data lake and analytical layers.

- Partner closely with Product, Underwriting, Analytics, and AI stakeholders to understand data requirements and translate them into robust data structures, contracts, and SLAs.

- Implement data quality controls, monitoring, and alerting to ensure accuracy, completeness, timeliness, and lineage of critical datasets and features used by models. Include automated controls for schema change, validity, reconciliation, claims maturity, and model leakage.

- Optimize serverless workloads for cost, performance, and scalability, including query tuning in Athena and efficient storage formats/partitioning in S3.

- Contribute to and enforce data engineering best practices, including version control, code review, CI/CD for data pipelines, and Infrastructure as Code with AWS CDK.

- Collaborate with AI and Analytics teams to design and maintain feature stores and other reusable data assets that accelerate experimentation and model deployment. Partner on batch or API scoring and capture model/data versions, recommendations, actions, overrides, and claims to close the learning loo

← All remote jobs