22 sep
|
Centraprise
|
México
22 sep
Centraprise
México
Job Title: Gen AI/ML Evaluation Engineer
Experience: 3–8 Years
Openings: 1–2
Focus: Gen AI/ML Evaluation, Data Pipelines, Test Infrastructure, NLP
Location: Mexico, Remote
Role Overview
We are looking for an engineer to build an automated evaluation and regression framework for a PII rewriting/sanitization service. The role involves building Golden Datasets, processing human annotations, developing evaluation pipelines, and measuring PII protection and rewrite quality.
Key Responsibilities
Build and maintain Golden Datasets from human annotations. Develop Python/data pipelines to extract and process data from Google Docs, Sheets, Slides, Gmail, Calendar, Chat, PDF, and Office files. Design data schemas/protobufs for machine-readable evaluation data. Build automated evaluation and regression test frameworks. Implement precision, recall, entity-level accuracy, mention-level accuracy, and rewrite consistency metrics. Develop automated reporting and regression detection. Support multilingual evaluation across English, Spanish, Portuguese, Japanese, and Korean. Use Gen AI/LLMs for evaluation,
annotation validation, error analysis, and quality measurement. Integrate evaluation frameworks with existing test/CI infrastructure.
Must-Have Skills
3+ years of software engineering experience with Python, C++, Java, or Go. Strong Python and data-processing experience. Experience building test automation, regression frameworks, or data pipelines. Hands-on experience with NLP, ML, Gen AI, or LLM-based systems. Strong understanding of precision, recall, false positives/negatives, and evaluation metrics. Experience with APIs, structured data, protobufs, and large datasets. Experience building or working with evaluation datasets/benchmarks. Strong debugging, data analysis, and problem-solving skills.
Preferred
Gen AI/LLM evaluation or benchmarking experience. PII detection, anonymization, or data privacy experience. Document parsing or Google Workspace API experience. Multilingual NLP experience. CI/CD and large-scale test infrastructure experience.
📌 Genai/ml evaluation engineer (México)
🏢 Centraprise
📍 México