Hello,
Greetings from ZettaMine!!!
Freelancing Opportunity :- We are currently hiring AI Benchmark Quality Reviewers for a short-term AI evaluation and benchmark quality project across LATAM.
Role: AI Benchmark Quality Reviewer
Experience: 5+ Years of Relevant Experience
Location: LATAM
Mode of Work: Remote
Engagement: Full-Time – 40 Hours/Week
Start Date: Immediate
PST Overlap: 8 Hours/Day
Description
We are seeking experienced AI Benchmark Quality Reviewers to validate the quality, correctness, and fairness of AI benchmark tasks and evaluation systems.
The role involves reviewing task instructions, reference solutions, grading criteria, agent execution traces, tool calls, and generated deliverables to identify task defects, grading issues, environment failures, and other quality concerns.
Responsibilities
- Validate task quality by checking instructions, source materials, reference solutions, and evaluation criteria for consistency.
- Identify hidden requirements, missing information, ambiguous instructions, or other task defects.
- Review agent performance by inspecting execution traces, tool calls, and generated deliverables.
- Determine whether agent successes and failures are appropriately justified.
- Audit grading logic and identify brittle checks, incorrect expected answers, unsupported rubric criteria, or unfair evaluation conditions.
- Identify cases where valid alternative solutions may be incorrectly penalized.
- Investigate discrepancies between model performance, task defects, grader errors,
and environment or tool failures.
- Independently assess automated QC findings rather than accepting them at face value.
- Document concise, evidence-backed findings and actionable feedback.
- Flag uncertainty and verify that revisions successfully resolve identified issues.
Requirements
- 5+ years of relevant professional experience.
- Comfortable reading Python, SQL, shell scripts, structured data, and execution logs to understand task setup and grading behavior.
- Strong analytical and problem-solving skills.
- Ability to independently verify calculations and reconcile conflicting evidence.
- Ability to assess the correctness and completeness of technical deliverables.
- Strong written English communication skills.
- Excellent attention to detail.
- Ability to provide specific, clear, and reproducible feedback.
- Comfortable working independently in a remote environment.
Preferred Experience
- AI evaluation or AI quality assurance.
- Technical QA or software quality review.
- Data analysis.
- AI benchmark development.
- Experience reviewing automated grading systems.
- Familiarity with Harbor task setup.
- Experience working with AI agents, execution traces, or tool-based workflows.
Availability
- Full-Time – 40 Hours per Week
- 8 Hours of PST overlap per day – Mandatory
- Immediate availability preferred.
Interested candidates kindly share your updated CV to
[email protected]
Thanks & Regards,
RaviTeja K
ZettaMine
📌 AI Benchmark Quality Reviewer (México)
🏢 ZettaMine Labs
📍 México