Senior Data Platform Engineer
Who we are
DigiCert is a global leader in intelligent trust. We protect the digital world by ensuring the security, privacy, and authenticity of every interaction. Our AI-powered DigiCert ONE platform unifies PKI, DNS, and certificate lifecycle management, to secure infrastructure, software, devices, messages, AI content and agents. Learn why more than 100,000 organizations, including 90% of the Fortune 500, choose DigiCert to stop todayβs threats and prepare for a quantum-safe future at
Job summary
We are looking for a Senior Data Platform Engineer to design, build, and operate the foundational data and ML infrastructure that powers analytics, reporting, and machine learning across DigiCert . This role sits at the intersection of data engineering and ML platform work β you will own the systems, pipelines, and tooling that data scientists, analysts, and engineers rely on every day. You bring deep expertise in Databricks, a strong engineering mindset, and hands-on experience building and maintaining ML pipelines in production. You will partner closely with data science, analytics, product, and business teams to deliver a platform that is reliable, governed, and built to scale.
What you will do
- Design, build, and maintain scalable data and ML pipelines using Python and SQL, processing large-scale datasets across batch and streaming workloads
- Own and evolve core platform infrastructure on Databricks β including Delta Lake table architecture, Unity Catalog governance, Databricks Workflows orchestration, and compute optimization
- Build and maintain end-to-end ML pipelines: feature engineering, model training pipelines, experiment tracking (MLflow), and model deployment/serving infrastructure
- Collaborate with data scientists to operationalize models β bridging the gap between experimentation and production-grade ML systems
- Define and enforce data platform standards: ingestion patterns, data modeling conventions, medallion architecture (Bronze/Silver/Gold), and pipeline reliability practices
- Implement data quality, observability, and monitoring frameworks to ensure platform health and data trustworthiness
- Optimize pipelines for performance, cost, and reliability at scale using Spark and PySpark
- Evaluate, integrate, and govern new platform tooling and data sources within the Databricks ecosystem
- Contribute to architectural decisions and help drive the long-term data platform roadmap
- Participate in code reviews, technical design discussions, and engineering standards
- Mentor junior engineers and elevate overall platform and data engineering practices
- Document platform architecture, pipeline design, and operational runbooks
What you will have
- 6+ years of experience in data engineering, data platform, or ML engineering roles
- Strong proficiency in Python and SQL, with a track record of building production-grade data pipelines using both
- Hands-on Databricks expertise: Delta Lake, Unity Catalog, Databricks Workflows, PySpark, and the Databricks ecosystem broadly
- Experience building and maintaining ML pipelines in production β feature engineering, training pipelines, experiment tracking, and model deployment
- Familiarity with MLflow or comparable experiment tracking and model registry tools
- Experience working on cloud data platforms (AWS, Azure, or GCP)
- Strong understanding of data modeling, dimensional design, and analytics-friendly data architecture
- Experience with batch and incremental/CDC pipeline patterns
- Proficiency with Git, version control, and CI/CD practices for data and ML workflows
- Strong engineering judgment β you think about reliability, maintainability, and cost, not just correctness
- Clear communication and comfort working with both technical and non-technical stakeholders
Nice to have
- Experience with streaming or near real-time pipelines (Kafka, Kinesis, Spark Structured Streaming)
- Familiarity with feature store platforms (Databricks Feature St