03 oct
|
eDataBae
|
Ciudad de México
03 oct
eDataBae
Ciudad de México
We are looking for experienced AI Benchmark Quality Reviewers to support an advanced AI evaluation project focused on improving the quality and reliability of AI-generated solutions. The role involves reviewing AI evaluation tasks, analyzing model outputs, validating grading logic, and providing evidence-based feedback to improve benchmark quality. ID 107627 _ AI Benchmark Qualit… What You'll Do Review task instructions, source materials, reference solutions, and evaluation criteria for accuracy and completeness Analyze AI agent execution traces, tool calls, and generated outputs Identify issues in grading logic, expected answers, and evaluation criteria Investigate whether failures are caused by model limitations, task issues, grader errors, or environment problems Provide clear,
structured feedback and document findings with supporting evidence ID 107627 _ AI Benchmark Qualit… Requirements 5+ years of relevant professional experience Comfortable reading and understanding: Python SQL Shell scripts Structured data Execution logs Strong analytical and problem-solving skills Ability to verify calculations, reconcile conflicting information, and assess technical deliverables Strong written English communication skills High attention to detail and ability to provide clear, reproducible feedback Skills: data analysis,technical quality assurance,ai evaluation #J-18808-Ljbffr
📌 Ai benchmark quality reviewer – ai evaluation project (latam - mexico) (Ciudad de México)
🏢 eDataBae
📍 Ciudad de México