Lead Data Engineer
Our vision for the future is based on the idea that transforming financial lives starts by giving our people the freedom to transform their own. We have a flexible work environment, and fluid career paths. We not only encourage but celebrate internal mobility. We also recognize the importance of purpose, well-being, and work-life balance. Within Empower and our communities, we work hard to create a welcoming and inclusive environment, and our associates dedicate thousands of hours to volunteering for causes that matter most to them.
Chart your own path and grow your career while helping more customers achieve financial freedom. Empower Yourself.
Applicants must be authorized to work for any employer in the U.S. We are unable to sponsor or take over sponsorship of an employment visa at this time, including CPT/OPT.
The Lead Data Engineer will provide technical leadership in the design, development, and optimization of scalable data pipelines, real-time ingestion frameworks, and data platforms. This role will lead complex data initiatives, including event-driven and streaming solutions, while ensuring performance, scalability, reliability, security, data quality, and governance align with enterprise standards. Operating with significant independence, the Lead Data Engineer will guide technical decision-making, mentor engineers, collaborate with architects and stakeholders, and serve as an escalation point for complex data engineering challenges.
What you will do:
-
Lead the design and implementation of advanced data pipelines, ingestion frameworks, and real-time streaming solutions.
-
Build, tune, and maintain scalable Kafka producers and consumers using Python, Java, Go, or Scala to support sub-second latency and high availability across distributed services.
-
Configure and manage enterprise change data capture (CDC) solutions, including Debezium, AWS DMS, or Qlik Replicate, across relational and NoSQL source systems such as PostgreSQL, MySQL, MongoDB, and Oracle into Kafka topics.
-
Develop stream-processing logic using Apache Flink, Spark Streaming, Kafka Streams, or dbt-mesh architectures to enrich, filter, and aggregate event payloads in flight.
-
Implement and enforce schema evolution policies using Confluent Schema Registry with Avro, Protobuf, or JSON Schema, and support appropriate delivery guarantees and idempotency across consumers.
-
Integrate automated code generation, validation agents, and LLM-driven orchestration using technologies such as LangGraph, AutoGen, or CrewAI to support schema drift, unstructured event parsing, and dynamic routing within event streams.
-
Instrument pipelines with end-to-end monitoring and alerting using tools such as Prometheus, Grafana, or Datadog, and manage consumer group offsets, partition rebalancing, and dead-letter queues.
-
Guide the technical execution of complex data initiatives across teams and translate highly complex business requirements into scalable data solutions.
-
Ensure data platforms are optimized for performance, scalability, reliability, and high availability.
-
Enforce data governance, security, quality, data contract, and schema management standards.
-
Mentor engineers, provide technical leadership on projects, and lead code reviews to ensure adherence to engineering standards.
-
Collaborate with architects and stakeholders to align solutions with enterprise direction and contribute to architectural and design decisions.
-
Drive improvements in automation, efficiency, and platform capabilities.
- Complete other duties as assigned.
What you will bring:
-
Bachelor’s degree in computer science or a related field, or an equivalent combination of training and experience.
-
6-8+ years of data engineering experience with enterprise scale platforms.
-
6-8+ years of hands-on experience with Apache Kafka, Confluent Cloud, AWS Kinesis, or Apache Pulsar.
-
Advanced proficiency in Python, PySpark, and SQL.
-
Proficiency in Python
Get remote data science jobs like this by email
One weekly digest. No spam, unsubscribe anytime.