Job Title: GenAI Evaluation & Data Infrastructure
Location: Mexico (Remote)
Experience: 8 years
Openings: 1–2 engineers
Focus: GenAI/ML Evaluation, Data Pipelines, Test Infrastructure, Workspace/Document Processing
Role Overview
We are looking for engineers to build an automated evaluation platform for a PII rewriting/sanitization service. The engineer will develop a machine-interpretable Golden Dataset from human annotations and build an automated regression framework to continuously measure rewrite quality, privacy protection, and consistency.
This role is particularly suited to engineers with experience in GenAI evaluation, NLP/ML systems, data pipelines, or large-scale test infrastructure.
Key Responsibilities
· Design and implement a structured Golden Dataset representing PII ground-truth annotations.
· Build tooling to extract annotations from human-rated evaluation artifacts.
· Develop parsers/extraction pipelines across:
· Google Docs
· Google Sheets
· Google Slides
· Gmail/Email
· Google Calendar
· Google Chat
· PDF and Microsoft Office files
· Design protobuf/data schemas for machine-interpretable evaluation data.
· Build an automated evaluation/regression harness for the PII sanitization/rewrite service.
· Implement evaluation metrics including:
· PII Recall
· PII Precision
· Rewrite Consistency
· Entity-level and mention-level accuracy
· Build automated reporting and regression detection for service quality.
· Extend evaluation coverage across English, Spanish, Portuguese, Japanese, and Korean.
· Investigate opportunities to use GenAI/LLMs for evaluation, annotation validation, error analysis, and quality measurement where appropriate.
· Integrate the evaluation framework into existing engineering/test infrastructure.
· Partner with ML, privacy, data, and Workspace engineers to validate evaluation methodology and dataset quality.
Required Qualifications
· Bachelor's degree or equivalent experience in Computer Science, Engineering,
or a related technical field.
· 3 years of software engineering experience with Python, C , Java, or Go.
· Strong experience building production-quality data processing or test infrastructure.
· Experience with structured data formats, APIs, protobufs, and large-scale datasets.
· Strong understanding of automated testing, regression testing, and evaluation frameworks.
· Experience working with NLP, ML, GenAI, LLMs, or text-processing systems.
· Ability to design reliable evaluation metrics and reason about precision/recall and false-positive/false-negative tradeoffs.
· Strong debugging, data analysis, and problem-solving skills.
Preferred Qualifications
· Experience building GenAI/LLM evaluation frameworks or benchmarks.
· Experience with PII detection, anonymization, data privacy, or synthetic data generation.
· Experience processing complex document formats or Workspace-like productivity artifacts.
· Experience with multilingual NLP or internationalization.
· Experience with entity/mention-level annotation and ground-truth dataset construction.
· Experience integrating automated evaluation into CI/CD or large-scale regression systems.
· Familiarity with Google Workspace APIs or similar document/email/calendar/chat systems.
Ideal Candidate Profile -
The adecuado candidate combines strong software engineering fundamentals with GenAI/ML evaluation experience. They should be comfortable moving between data modeling, document parsing, evaluation methodology, test infrastructure, and ML/LLM-based quality analysis.
Resource Split
GenAI/ML Evaluation Engineer
· Golden dataset design
· Annotation schema
· Evaluation methodology
· Precision/recall/consistency metrics
· GenAI-assisted evaluation and error analysis
· Multilingual evaluation
Software/Data Infrastructure Engineer
· Artifact extraction
· Workspace/document integrations
· Proto/data pipelines
· Dataset generation tooling
· Regression test harness
Automated reporting and CI integration
Kindly share the resume to
[email protected]
📌 GenAI Evaluation & Data Infrastructure (México)
🏢 Ampstek
📍 México