Senior Data Engineer Consultant
Job title: Senior Data Engineering Consultant -Platform Architecture & AI-Native Data Strategy
Engagement type: Contract / Consulting (3โ6 months with ongoing advisory)
Location: Remote (reasonable 3-4 hrs overlap with US working hours required)
Experience: 8-10+ years in data engineering and data platform architecture
About DiligenceVault
DiligenceVault is an enterprise B2B SaaS platform that helps institutional investors, asset managers, consultants, and fund service providers digitize and automate the end-to-end due diligence lifecycle. The platform supports workflows including DDQs, RFPs, Operational Due Diligence (ODD), manager research, ESG, compliance, and investor reporting through AI-powered document processing, workflow automation, analytics, and collaboration. Today, DiligenceVault serves 100,000+ platform users, 20,000+ managers, and 250+ client teams across 150+ countries.
Our platform processes large volumes of structured and unstructured data from customer-uploaded documents, digital questionnaires, CRM systems, enterprise content repositories, regulatory filings, and platform-generated workflow data. Our current technology stack includes Azure SQL/SQL Server, Python/Celery, .NET REST APIs, Elasticsearch, Azure OpenAI, and Kestra for orchestration.
We want to build a deliberate data platform that turns this raw data into meaningful customer intelligence. We need a consultant who can help us understand the full landscape of data engineering (traditional and AI-native), assess where we are, and architect where we need to go.
What you will do
Phase 1- Educate and assess
Teach our leadership and senior architects the full spectrum of data engineering, covering traditional foundations and AI-native approaches in depth. This is not a surface-level overview - our team needs to understand concepts deeply enough to make architectural decisions. Topics span ingestion patterns (batch, streaming, CDC, adaptive connectors), transformation (ETL/ELT, dbt, Spark, LLM-assisted mapping), data modeling (dimensional, data vault, lakehouse, schema-on-read), data quality (rule-based vs. ML-driven anomaly detection, data contracts), entity resolution and data stitching (manual mapping vs. embedding-based semantic matching, knowledge graphs), orchestration (DAG engines, event-driven, self-healing pipelines), semantic layers (ontologies, contextual meaning, embedding-based search), and AI-native versioning (prompts, models, thresholds, reproducibility).
Assess our current data infrastructure end to end. Map existing data flows, identify gaps and technical debt, and produce a landscape assessment with current state, target state, gap analysis, and a prioritized roadmap.
Phase 2 - Architect and define use cases
Design the target data platform architecture across ingestion, transformation, storage, serving, and observability layers. Within this architecture, four strategic initiatives require specific attention:
PostgreSQL migration and multi-workload architecture. We are planning to move from SQL Server to PostgreSQL. The consultant will help architect a PostgreSQL environment that supports multiple workload types through the PostgreSQL extension ecosystem - pgvector for vector similarity search and embedding storage powering our AI features, analytical query patterns (columnar extensions like Citus or pg_analytics, or appropriate separation of OLAP workloads), and transactional queries for the core application. This includes guidance on connection pooling (PgBouncer/PgCat), read replica topology, partitioning strategies, and how to handle workloads that on SQL Server relied on specific features (stored procedures, Query Store, tempdb patterns) that work differently in PostgreSQL. The migration path itself - phased cutover strategy, dual-write/shadow-read validation, query translation, and performance benchmarking - is a key deliverable.
Canonical data architecture across heterogeneous sources. Data arrives from dozens of