Forward Deployed Engineer (Data, ML & AI)
Forward Deployed Engineer (Data, ML & AI)
Location: Remote Timezone: NA/Eastern Type: Full-time Experience: 10+ Years
Role Overview
This position requires a Forward Deployed Engineer (FDE) specializing in Data, Machine Learning, and AI to embed directly within customer environments. You will serve as the primary technical authority transforming complex data challenges and operational bottlenecks into production-grade data pipelines, machine learning systems, and agentic AI solutions.
The role is heavily customer-facing: you will work alongside client business units and engineering teams to rapidly assess legacy data estates, build scalable modern data platforms, and deploy custom GenAI/LLM workflows that drive measurable business velocity. This is a high-ownership, hands-on role where you will leverage Spec-Driven Development (SDD) and AI-assisted workflows to build, refactor, and demonstrate immediate value directly where the customer operates.
Key Responsibilities
- Customer Embedding & Data Delivery: Deploy directly into customer environments to understand their domain logic, underlying data pipelines, and architectural constraints. Own end-to-end delivery—from data discovery and schema design to pipeline deployment and model integration—building trust as the customer's lead technical partner.
- Data Platform & Architectural Modernization: Partner with client business leaders to evaluate legacy technology estates (e.g., monolithic SQL databases, or unmaintained ETL jobs). Lead engineering efforts to refactor legacy data setups into modern lakehouses, event-driven streaming systems, and scalable vector/graph databases.
- Spec-Driven Development (SDD) for Data & AI: Apply a "spec-first" engineering workflow using Generative AI tools. Write structured specifications (data models, API schemas, transformations, and evaluation metrics) that instruct AI agents to generate production data models, PySpark jobs, data pipelines, and test suites.
- Agentic AI & LLMOps Implementation: Architect and deploy GenAI workflows, Retrieval-Augmented Generation (RAG) pipelines, and autonomous AI agents using frameworks such as LangChain, LlamaIndex, or DSPy. Establish robust evaluation frameworks (Evals) for model accuracy, latency, and hallucination control.
- Production MLOps & Orchestration: Build, deploy, and maintain robust ML training and inference pipelines using tools like MLflow, Kubeflow, Airflow, or Dagster. Ensure continuous integration/continuous deployment (CI/CD) for models and data workflows.
- Polyglot Data Engineering: Design and audit production code across data-centric languages and frameworks (Python, SQL, Scala, Go, Rust, or TypeScript) based on speed, concurrency, and memory requirements.
- Client Enablement & Knowledge Transfer: Elevate customer teams by establishing reusable agentic development patterns, modern MLOps practices, data reliability frameworks, and SDD methodologies so systems remain maintainable long after deployment.
Requirements
- 10+ Years of Experience: Proven track record as a Principal Data Engineer, Lead ML Engineer, or Enterprise Data Architect building and scaling distributed data and ML platforms.
- Customer-Facing Aptitude: Strong executive presence and communication skills to interface directly with technical teams and business stakeholders under pressure.
- Data & ML Engineering Depth:
- Data Infrastructure: Mastery of distributed computing (AWS Glue, Apache Spark, Databricks), modern data warehouses (Redshift, Snowflake, modeling tools (dbt), and data orchestration (Airflow, etc)
- AI/ML & Vector Architecture: Hands-on experience fine-tuning, evaluating, and deploying LLMs, embedding models, and vector stores
- Polyglot & Framework Proficiency: Advanced proficiency in Python and complex SQL, plus fluency in at least two other languages used in modern backend/data systems (e.g., Scala, Go, Rust, TypeScript).
- Generative AI & SDD Experience: Demonstrated skill in using natur