Staff Engineer - Senior Data Modeler
Role - Senior Data Modeler
Working Model - 100% Remote (USA/Canada)
Employment type โ Fulltime
Numbers of Interviews - 2 or 3
Role Summary:
- Senior Data Modeler Unstructured Data / Knowledge Management
- The Senior Data Modeler will design and govern the data architecture for unstructured knowledge assets across Client's Knowledge Management (KM) ecosystem.
- This role bridges data engineering discipline with KM domain expertise, translating raw unstructured content (documents, case files, informal knowledge captures, chat/email extracts, etc.) into well-defined, discoverable, and secure data products within Databricks Unity Catalog.
- This is a foundational hire for a newly forming KM Data Platform team supporting Client's broader Knowledge and Research Systems strategy.
Mandatory Skills:
- 7+ years of experience in data modeling, information architecture or enterprise data architecture.
- Strong experience designing conceptual, logical and physical data models for enterprise data platforms.
- Strong understanding of entities, relationships, metadata, master/reference data and data lineage.
- Experience with taxonomy, ontology, semantic models and controlled vocabularies.
- Hands-on experience with Databricks, Delta Lake and Unity Catalog or comparable modern data platforms.
- Ability to translate business concepts and unstructured information into structured, reusable data models.
Good to Have Skills:
- Experience designing knowledge graphs, ontologies and semantic knowledge models.
- Experience with GenAI/RAG knowledge models and vector/embedding representations.
- Experience modeling documents, document elements, entities, relationships, evidence and provenance.
- Knowledge of knowledge graph technologies such as Neo4j, RDF or property graphs.
- Experience with Commercial/Customer/CRM domain models, SharePoint content or enterprise knowledge platforms.
Key Responsibilities:
- Design logical and physical data models for unstructured and semi-structured content (documents, case artifacts, K-Slices, extracted knowledge fragments, metadata records) originating from KM pipelines such as case mining and informal knowledge capture workflows.
- Define domain boundaries and ownership for data products โ determining what constitutes a discrete, reusable data product versus a raw or intermediate asset.
- Establish metadata standards and tagging taxonomies (content type, practice/domain, provenance, confidentiality, freshness, lineage) to ensure consistent classification across knowledge sources.
- Assign and enforce security and sensitivity classifications on data products in line with firm data governance, privacy, and legal/risk requirements.
- Register, document, and maintain data products in Databricks Unity Catalog, including schemas, access grants, lineage, and catalog-level metadata.
- Partner with data engineers building Databricks pipelines to ensure ingestion, transformation, and storage patterns align to the modeled domain structure.
- Collaborate with Knowledge Products, Research Products, and Architecture/Data/Technology stakeholders to align data product design with downstream consumption needs (e.g., surfacing in Sage/Glean, AI agent retrieval).
- Support privacy and legal review processes by ensuring data products are classified and documented to enable timely sign-off.
- Establish and document repeatable modeling standards/playbooks so future data products can be onboarded consistently as the KM platform scales.
Required Qualifications:
- 5+ years of experience in data modeling, data architecture, or information architecture, with meaningful exposure to unstructured or semi-structured data (not purely relational/transactional modeling).
- Direct experience working in or adjacent to Knowledge Management, content management, or enterprise search domain โ understands how documents, case files, or knowledge artifacts differ from standard transactional data.
- Hands-on
Get remote data science jobs like this by email
One weekly digest. No spam, unsubscribe anytime.