We are seeking a Data Scientist to support Conversational Analytics and Generative AI initiatives through the development of scalable data pipelines, prompt engineering approaches, sampling methodologies, and Athena Studio workflows for large-scale conversational data. This role will collaborate closely with Data Science, Engineering, and Product teams to analyze conversational datasets, drive AI-powered insights, and help build adaptive technologies that enable action across the organization. Location:Remote, Mexico only. About Us: Abstra is a fast-growing, Nearshore Tech Talent services company, providing top Latin American tech talent to U. S. companies and beyond. Founded by U. S. -bred engineers with over 15 years of experience, Abstra specializes in sourcing skilled professionals across a wide range of technologies to meet our clients’ needs, driving innovation and efficiency. Responsibilities: - As a Data Scientist on the Data Science team, you will develop machine learning, NLP and Generative AI solutions focused on conversational analytics and large language models. - Design, develop, test and optimise prompts for Large Language Models (LLMs) using conversational datasets. - Support prompt training and evaluation activities to improve accuracy, relevance and consistency of AI-driven outputs. - Create scalable data pipelines for ingestion, transformation, sampling and processing of conversational and unstructured text data. - Develop sampling methodologies for high-volume conversational datasets, including sampling within conversations and across conversations. - Prepare and process conversational data for use in Athena Studio and related analytics workflows. - Apply NLP and machine learning techniques for text understanding, classification, summarization, pattern recognition and conversational insights.
- Run data science analyses and experiments independently while collaborating with engineering, analytics and product teams.
- Communicate technical methods, results and recommendations clearly to both technical and non-technical stakeholders. Minimum Qualifications - Bachelor’s degree in Computer Science, Data Science, Artificial Intelligence, Machine Learning, Statistics, Mathematics or related technical discipline, or equivalent practical experience. - 5+ years of relevant work experience in Data Science, Machine Learning, Artificial Intelligence, NLP or analytics engineering. - Hands-on experience with Prompt Engineering and Large Language Models (LLMs). - Experience creating data pipelines, ETL/ELT workflows or data processing pipelines for structured and unstructured datasets. - Experience working with conversational data, text analytics, customer interaction data or other large-scale unstructured data sources. - Demonstrated proficiency in Python and SQL. - Working knowledge of packages and frameworks associated with data wrangling, analysis and machine learning, such as Pandas, NumPy, Scikit-learn, SpaCy, TensorFlow, PyTorch, LangChain or LangGraph. Preferred Qualifications - MS degree in Computer Science, Artificial Intelligence, Machine Learning, Data Science or related technical discipline. - Experience using Athena Studio and AWS Athena for analytics or data processing workflows. - Experience with prompt training, prompt evaluation, LLM evaluation frameworks or fine-tuning concepts. - Experience in one or more of the following areas: Natural Language Processing for text understanding, classification, summarization, recommendation/ranking systems or conversational analytics. - Experience with cloud data and machine learning platforms such as AWS S3, AWS Glue, SageMaker, Databricks, Vertex AI or Kubeflow. - Experience with embeddings, semantic search, vector databases, RAG patterns or representation learning for unstructured data. - Strong communication skills, especially in describing technical methods and results to both technical and non-technical audiences.
📌 Data Scientist (Generative AI, LLM & Conversational Analytics) (México)
🏢 Abstra
📍 México