Data Scientist (Generative AI, LLM & Conversational Analytics)

🏢 Abstra · all Abstra jobs
📍 Mexico
📅 Posted 2026-08-22 · via Himalayas
🏷 Data-Scientist,Generative-AI,LLM-Engineering,Conversational-Analytics,NLP,Prompt-Engineering,AI-LLM-Data-Scientist,AI-Data-Scientist,NLP-Data-Scientist,Remote-Senior-Generative-AI-Scientist,ML-Data-Scientist,Generative-AI-Engineer
Apply on original site ↗

We are seeking a Data Scientist to support Conversational Analytics and Generative AI initiatives through the development of scalable data pipelines, prompt engineering approaches, sampling methodologies, and Athena Studio workflows for large-scale conversational data. This role will collaborate closely with Data Science, Engineering, and Product teams to analyze conversational datasets, drive AI-powered insights, and help build adaptive technologies that enable action across the organization.

Location:Remote, Mexico only.

About Us:

Abstra is a fast-growing, Nearshore Tech Talent services company, providing top Latin American tech talent to U.S. companies and beyond. Founded by U.S.-bred engineers with over 15 years of experience, Abstra specializes in sourcing skilled professionals across a wide range of technologies to meet our clients’ needs, driving innovation and efficiency.

Responsibilities:

- As a Data Scientist on the Data Science team, you will develop machine learning, NLP and Generative AI solutions focused on conversational analytics and large language models.

- Design, develop, test and optimise prompts for Large Language Models (LLMs) using conversational datasets.

- Support prompt training and evaluation activities to improve accuracy, relevance and consistency of AI-driven outputs.

- Create scalable data pipelines for ingestion, transformation, sampling and processing of conversational and unstructured text data.

- Develop sampling methodologies for high-volume conversational datasets, including sampling within conversations and across conversations.

- Prepare and process conversational data for use in Athena Studio and related analytics workflows.

- Apply NLP and machine learning techniques for text understanding, classification, summarization, pattern recognition and conversational insights.

- Run data science analyses and experiments independently while collaborating with engineering, analytics and product teams.

- Communicate technical methods, results and recommendations clearly to both technical and non-technical stakeholders.

Minimum Qualifications

- Bachelor’s degree in Computer Science, Data Science, Artificial Intelligence, Machine Learning, Statistics, Mathematics or related technical discipline, or equivalent practical experience.

- 5+ years of relevant work experience in Data Science, Machine Learning, Artificial Intelligence, NLP or analytics engineering.

- Hands-on experience with Prompt Engineering and Large Language Models (LLMs).

- Experience creating data pipelines, ETL/ELT workflows or data processing pipelines for structured and unstructured datasets.

- Experience working with conversational data, text analytics, customer interaction data or other large-scale unstructured data sources.

- Demonstrated proficiency in Python and SQL.

- Working knowledge of packages and frameworks associated with data wrangling, analysis and machine learning, such as Pandas, NumPy, Scikit-learn, SpaCy, TensorFlow, PyTorch, LangChain or LangGraph.

Preferred Qualifications

- MS degree in Computer Science, Artificial Intelligence, Machine Learning, Data Science or related technical discipline.

- Experience using Athena Studio and AWS Athena for analytics or data processing workflows.

- Experience with prompt training, prompt evaluation, LLM evaluation frameworks or fine-tuning concepts.

- Experience in one or more of the following areas: Natural Language Processing for text understanding, classification, summarization, recommendation/ranking systems or conversational analytics.

- Experience with cloud data and machine learning platforms such as AWS S3, AWS Glue, SageMaker, Databricks, Vertex AI or Kubeflow.

- Experience with embeddings, semantic search, vector databases, RAG patterns or representation learning for unstructured data.

- Strong communication skills, especially in describing technical methods and results to both technical and non-technical audiences.

- Ability to collab

← All remote jobs