8-12 Weeks
- 8+ yrs of relevant Experience
- Comfortable reading Python, SQL, shell scripts, structured data, and execution logs to understand task setup and grading behavior.
Responsibilities:
- Validate task quality: Check that instructions, source materials, reference solutions, and evaluation criteria are consistent, with no hidden requirements or missing information.
- Review agent performance: Inspect execution traces, tool calls, and generated deliverables to determine whether successes and failures are justified.
- Audit grading logic: Identify brittle checks, incorrect expected answers, unsupported rubric criteria, and cases where valid alternative solutions are unfairly penalized.
- Investigate discrepancies: Distinguish genuine model limitations from task defects, grader errors, and environment or tool failures. Assess automated QC findings independently rather than accepting them at face value.
- Document decisions: Provide concise,
evidence-backed findings and actionable feedback, flag uncertainty, and verify that revisions resolve identified issues.
What we're looking for:
- Technical fluency: Comfortable reading Python, SQL, shell scripts, structured data, and execution logs to understand task setup and grading behavior.
- Analytical judgment: Able to independently check calculations, reconcile conflicting evidence, and assess the correctness and completeness of professional deliverables.
- Clear communication: Strong written English, attention to detail, and experience providing specific, reproducible feedback.
- Relevant experience preferred: AI evaluation, technical QA, data analysis, or benchmark development; familiarity with Harbor task setup.
📌 AI Benchmark Quality Reviewer (México)
🏢 Turing
📍 México