Role: AI Training Data Quality Reviewer – Python/SQL
Location: LATAM (Remote)
Term: Conract
Role Overview
Looking for experienced professionals to join as AI Benchmark Quality Reviewers — auditing the quality of AI evaluation tasks and grading pipelines that power frontier AI model development.
What you'll do:
- Validate AI task quality — instructions, source materials, reference solutions & evaluation criteria
- Review AI agent execution traces, tool calls & deliverables
- Audit grading logic and flag brittle or unfair rubric checks
- Investigate discrepancies between model limitations and task/grader defects
- Document clear, evidence-backed findings
What we're looking for:
- 5+ years of relevant experience
- Comfortable reading Python, SQL, shell scripts & execution logs
- Strong analytical judgment and written English
- Bonus: background in AI evaluation, technical QA, or benchmark development (Harbor experience a plus
- If you love digging into logic, spotting inconsistencies, and shaping how AI systems are evaluated — this one's for you.
📌 AI Training Data Quality Reviewer - Python/SQL (Ciudad de México)
🏢 Epsilon Solutions
📍 Ciudad de México
Postulate a este anuncio
Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.