Applied AI Engineer - India

🏢 Dscout · all Dscout jobs
📍 India
📅 Posted 2026-08-30 · via Himalayas
🏷 Applied-AI-Engineer,AI-Product-Engineering,Machine-Learning-Engineering,LLM-Engineering,AI-Systems-Development,Applied-AI-Engineer-Jobs,AI-Application-Engineer,Applied-AI-ML-Engineer,AI-Applications-Engineer,Applied-AI-Engineering,Applied-Machine-Learning-Engineer,AI-Applications-Engineering,AI-Engineer
Apply on original site ↗

At Dscout , we’re building the most flexible and powerful UX research platform on the market—trusted by the world’s top brands in finance (JP Morgan Chase, Intuit, Charles Schwab, PayPal), healthcare (Aya, Headspace), consumer goods (Keen, Verizon, Target, Northface), and tech (Google, Amazon, Facebook, Meta, Spotify, AirBnB). Our tools help teams deeply understand the humans behind their products, so they can build better ones. We are expanding our smart and driven team and would love for you to join us.

AI native product is fundamentally different engineering problem than building deterministic software: the same input won't always produce the same output, and "working" means the agent behaves well across the full distribution of real world scenarios, not that it passes a fixed test suite.

We're looking for an Applied AI Engineer with 2-5 years of experience building and shipping AI systems used by professionals at enterprise. You're comfortable working with modern LLM-based systems and agentic workflows, and you know how to turn powerful models into reliable product features. You have strong product judgment and think deeply about tradeoffs between LLM approaches and traditional ML when designing solutions. You care about evaluation, iteration speed, and making sure AI systems actually drive measurable business impact reliably .
What you'll do

- Own the production improvement loop across agent behavior, customer and operator feedback, evaluation, experimentation, and verified business outcomes

- Instrument agent workflows so model interactions, tool use, decisions, failures, human edits, and downstream outcomes can be understood in context

- Define meaningful quality standards, representative evaluation datasets, regression coverage, and production monitoring.

- Investigate why agents underperform across context, knowledge, instructions, tools, routing, guardrails, or workflow design

- Design and ship targeted behavior improvements, including changes to prompting, context construction, decision logic, tool use, and human-review paths

- Build backend services, APIs, data models, and feedback pipelines that make agent behavior observable, steerable, and reproducible

- Run controlled experiments, production replays, or staged rollouts to measure whether changes improve quality and downstream business results

- Partner with Product, Data Science, and Sales to prioritize high-value problems and define customer and business success

- Ship with appropriate safeguards for privacy, security, reliability, human oversight, and safe operational rollout

What you bring

- 2-5 years of software engineering experience, with hands-on experience building or operating LLM-powered features or agents in production - not just prototypes or demos

- Fluency with prompting and context engineering as an engineering discipline: you iterate on prompts, context construction, and tool definitions the way you'd iterate on code

- Experience building or maintaining evaluation harnesses for AI systems: offline eval sets, LLM-as-judge or human-in-the-loop scoring, regression detection

- Genuine comfort with non-determinism: you reason about agent behavior across a distribution of production traffic, not a fixed set of test cases, and you don't treat variance as a bug to be argued away

- Experience running experiments (A/B, staged rollouts, production replay) to validate whether a change actually improved outcomes, not just whether it shipped

- A track record of shipping features real users depended on, and owning what happened after launch

- A high-agency mindset: comfortable investigating an ambiguous "why is this underperforming" problem across context, tools, routing, and workflow design without a fully-scoped ticket

- Comfort using AI coding tools (Cursor, Claude Code, Copilot, or similar) as a real part of your workflow

Nice to have

- Experience with voice or real-time conversational AI systems

- Familiarity with LLM observabili

← All remote jobs

Similar for you