02 oct
|
eDataBae
|
Ciudad de México
02 oct
eDataBae
Ciudad de México
We are looking for experienced AI Benchmark Quality Reviewers to support an advanced AI evaluation project focused on improving the quality and reliability of AI-generated solutions.
The role involves reviewing AI evaluation tasks, analyzing model outputs, validating grading logic, and providing evidence-based feedback to improve benchmark quality. ID 107627 _ AI Benchmark Qualit…
What You'll Do
Review task instructions, source materials, reference solutions, and evaluation criteria for accuracy and completeness
Analyze AI agent execution traces, tool calls, and generated outputsIdentify issues in grading logic, expected answers, and evaluation criteria
Investigate whether failures are caused by model limitations, task issues, grader errors, or environment problems
Provide clear,
structured feedback and document findings with supporting evidence ID 107627 _ AI Benchmark Qualit…
Requirements
5+ years of relevant professional experience
Comfortable reading and understanding:
Python
SQL
Shell scripts
Structured data
Execution logs
Strong analytical and problem-solving skills
Ability to verify calculations, reconcile conflicting information, and assess technical deliverables
Strong written English communication skills
High attention to detail and ability to provide clear, reproducible feedback
Skills: data analysis,technical quality assurance,ai evaluation
#J-18808-Ljbffr
📌 AI Benchmark Quality Reviewer – AI Evaluation Project (LATAM - Mexico) (Ciudad de México)
🏢 eDataBae
📍 Ciudad de México