Data Scientist (Generative AI, LLM & Conversational Analytics)
We are seeking a Data Scientist to support Conversational Analytics and Generative AI initiatives through the development of scalable data pipelines, prompt engineering approaches, sampling methodologies, and Athena Studio workflows for large-scale conversational data. This role will collaborate closely with Data Science, Engineering, and Product teams to analyze conversational datasets, drive AI-powered insights, and help build adaptive technologies that enable action across the organization.
Location:Remote, Mexico only.
About Us:
Abstra is a fast-growing, Nearshore Tech Talent services company, providing top Latin American tech talent to U.S. companies and beyond. Founded by U.S.-bred engineers with over 15 years of experience, Abstra specializes in sourcing skilled professionals across a wide range of technologies to meet our clients’ needs, driving innovation and efficiency.
Responsibilities:
- As a Data Scientist on the Data Science team, you will develop machine learning, NLP and Generative AI solutions focused on conversational analytics and large language models.
- Design, develop, test and optimise prompts for Large Language Models (LLMs) using conversational datasets.
- Support prompt training and evaluation activities to improve accuracy, relevance and consistency of AI-driven outputs.
- Create scalable data pipelines for ingestion, transformation, sampling and processing of conversational and unstructured text data.
- Develop sampling methodologies for high-volume conversational datasets, including sampling within conversations and across conversations.
- Prepare and process conversational data for use in Athena Studio and related analytics workflows.
- Apply NLP and machine learning techniques for text understanding, classification, summarization, pattern recognition and conversational insights.
- Run data science analyses and experiments independently while collaborating with engineering, analytics and product teams.
- Communicate technical methods, results and recommendations clearly to both technical and non-technical stakeholders.
Minimum Qualifications
- Bachelor’s degree in Computer Science, Data Science, Artificial Intelligence, Machine Learning, Statistics, Mathematics or related technical discipline, or equivalent practical experience.
- 5+ years of relevant work experience in Data Science, Machine Learning, Artificial Intelligence, NLP or analytics engineering.
- Hands-on experience with Prompt Engineering and Large Language Models (LLMs).
- Experience creating data pipelines, ETL/ELT workflows or data processing pipelines for structured and unstructured datasets.
- Experience working with conversational data, text analytics, customer interaction data or other large-scale unstructured data sources.
- Demonstrated proficiency in Python and SQL.
- Working knowledge of packages and frameworks associated with data wrangling, analysis and machine learning, such as Pandas, NumPy, Scikit-learn, SpaCy, TensorFlow, PyTorch, LangChain or LangGraph.
Preferred Qualifications
- MS degree in Computer Science, Artificial Intelligence, Machine Learning, Data Science or related technical discipline.
- Experience using Athena Studio and AWS Athena for analytics or data processing workflows.
- Experience with prompt training, prompt evaluation, LLM evaluation frameworks or fine-tuning concepts.
- Experience in one or more of the following areas: Natural Language Processing for text understanding, classification, summarization, recommendation/ranking systems or conversational analytics.
- Experience with cloud data and machine learning platforms such as AWS S3, AWS Glue, SageMaker, Databricks, Vertex AI or Kubeflow.
- Experience with embeddings, semantic search, vector databases, RAG patterns or representation learning for unstructured data.
- Strong communication skills, especially in describing technical methods and results to both technical and non-technical audiences.
- Ability to collab