Architect - Machine Learning and Data
Coretek is seeking an Architect, Machine Learning and Data to lead the design of production machine learning and generative AI platforms on Microsoft Azure and Microsoft Fabric. This role sits at the point where data platform architecture, MLOps, and applied AI meet. You will define how client organizations move models out of notebooks and into governed, monitored, reproducible production operation, and how generative AI capabilities are architected to be secure, evaluable, and cost-controlled at enterprise scale.
You will own the technical architecture on client engagements end to end: target-state design, environment and identity models, pipeline and promotion patterns, observability standards, and the operational handoff that lets a client run the platform without you.
You will work alongside data scientists, data engineers, and delivery leadership, translating requirements into architecture decisions and then staying close enough to implementation to be accountable for the result.
This is a delivery architecture role. Depth of production experience matters more than breadth of exposure, and the expectation is that you have personally been responsible for systems that ran unattended and were handed to someone else to operate.
Key Responsibilities:
Solution Architecture and Technical Strategy
- Own end-to-end technical architecture for machine learning and AI engagements, from target-state design through production acceptance.
- Define reference architectures for Fabric-first data science and MLOps platforms, including environment topology, storage boundaries, and promotion paths.
- Produce architecture decision records, requirements traceability, and design documentation that hold up under client security and compliance review.
- Make and defend platform tradeoff decisions: Fabric versus Azure-native services, managed versus custom components, build versus configure.
- Define compute sizing assumptions, cost guardrails, and capacity planning for batch and inference workloads.
- Establish third-party and open-source governance patterns, including dependency disclosure, licensing implications, and controls that keep unapproved packages out of production.
ML Platform, MLOps, and Operationalization
- Architect Sandbox, Dev/Staging, and Production environment models with enforced isolation and role-based access aligned to Entra ID group structures.
- Design governed read access to enterprise data warehouse sources alongside controlled data science owned write-back boundaries for features, model metadata, artifact references, predictions, and experiment results.
- Define reusable batch prediction and forecasting pipeline architectures spanning ingestion, feature preparation, quality validation, model execution, output persistence, and alerting.
- Architect forecasting-specific patterns where they diverge from batch scoring, including time-series inputs, rolling forecasts, and horizon-based outputs.
- Design CI/CD and promotion architecture for notebooks and platform assets: Git integration, branching standards, automated testing, deployment pipelines, approval gates, and rollback paths.
- Define orchestration and scheduling patterns covering time-based, trigger-based, and manual execution with dependency-level failure visibility.
- Architect data quality gates that block downstream model execution on failure, covering schema validation, null and range thresholds, and distributional anomaly detection.
- Mandate and design headless execution: all scheduled and production workloads run under managed identities or service principals with secrets in Azure Key Vault, never under individual user credentials.
- Establish model, code, environment, and package versioning standards so any production run is traceable to a versioned combination of code, configuration, environment, and data reference.
- Design observability and drift monitoring architecture, including baseline statistics, health checks, alert threshold