21 sep
|
Centraprise
|
Ciudad de México
21 sep
Centraprise
Ciudad de México
Job Title: Gen AI/ML Evaluation Engineer Experience: 3–8 Years Openings: 1–2 Focus: Gen AI/ML Evaluation, Data Pipelines, Test Infrastructure, NLP Location: Mexico, Remote Role Overview We are looking for an engineer to build an automated evaluation and regression framework for a PII rewriting/sanitization service. The role involves building Golden Datasets, processing human annotations, developing evaluation pipelines, and measuring PII protection and rewrite quality. Key Responsibilities Build and maintain Golden Datasets from human annotations. Develop Python/data pipelines to extract and process data from Google Docs, Sheets, Slides, Gmail, Calendar, Chat, PDF, and Office files. Design data schemas/protobufs for machine-readable evaluation data. Build automated evaluation and regression test frameworks. Implement precision, recall, entity-level accuracy, mention-level accuracy, and rewrite consistency metrics. Develop automated reporting and regression detection. Support multilingual evaluation across English, Spanish, Portuguese, Japanese, and Korean. Use Gen AI/LLMs for evaluation, annotation validation,
error analysis, and quality measurement. Integrate evaluation frameworks with existing test/CI infrastructure. Must-Have Skills 3+ years of software engineering experience with Python, C++, Java, or Go. Strong Python and data-processing experience. Experience building test automation, regression frameworks, or data pipelines. Hands-on experience with NLP, ML, Gen AI, or LLM-based systems. Strong understanding of precision, recall, false positives/negatives, and evaluation metrics. Experience with APIs, structured data, protobufs, and large datasets. Experience building or working with evaluation datasets/benchmarks. Strong debugging, data analysis, and problem-solving skills. Preferred Gen AI/LLM evaluation or benchmarking experience. PII detection, anonymization, or data privacy experience. Document parsing or Google Workspace API experience. Multilingual NLP experience. CI/CD and large-scale test infrastructure experience.
📌 Genai/ml evaluation engineer (Ciudad de México)
🏢 Centraprise
📍 Ciudad de México