08 oct
|
eDataBae
|
México
We are looking for experienced
AI Benchmark Quality Reviewers
to support an advanced AI evaluation project focused on improving the quality and reliability of AI-generated solutions.
The role involves reviewing AI evaluation tasks, analyzing model outputs, validating grading logic, and providing evidence-based feedback to improve benchmark quality. ID ****** _ AI Benchmark Qualit...
What You'll Do
Review task instructions, source materials, reference solutions, and evaluation criteria for accuracy and completeness
Analyze AI agent execution traces, tool calls, and generated outputs
Identify issues in grading logic, expected answers, and evaluation criteria
Investigate whether failures are caused by model limitations, task issues, grader errors, or environment problems
Provide clear,
structured feedback and document findings with supporting evidence ID ****** _ AI Benchmark Qualit...
Requirements
5+ years of relevant professional experience
Comfortable reading and understanding:
Python
SQL
Shell scripts
Structured data
Execution logs
Strong analytical and problem-solving skills
Ability to verify calculations, reconcile conflicting information, and assess technical deliverables
Strong written English communication skills
High attention to detail and ability to provide clear, reproducible feedback
📌 Ai Benchmark Quality Reviewer - Ai Evaluation Project (Latam - Mexico) (México)
🏢 eDataBae
📍 México