10 ago
|
Cafeto Software
|
México
10 ago
Cafeto Software
México
We're looking for someone to own the evaluation strategy for AI-powered workflows that handle real billing and collections outcomes — agent behavior, output quality, accuracy, and business-rule compliance.
What You'll Do
Build automated AI evaluations (DeepEval, Ragas, Promptfoo, LangSmith, or similar) combining deterministic checks, LLM-as-a-judge, and production evidence.
Create and maintain evaluation datasets: golden examples, regression suites, edge cases.
Integrate evaluations into CI/CD so AI regressions get caught before release.
Test the full workflow — APIs, backend services, async processing, third-party integrations — not just the model.
Investigate failures using Langfuse, CloudWatch, logs, and traces to separate product defects from model behavior.
Build targeted Playwright tests where browser-level coverage protects a critical workflow.
Give clear release-readiness input based on evidence, not gut feel.
What You Need
5+ years hands-on QE, SDET, or AI Quality experience.
Real experience testing LLM/AI/RAG/agentic systems — hallucinations, partial correctness, non-determinism.
Hands-on building AI evaluations (DeepEval strongly preferred; Ragas/Promptfoo/LangSmith/TruLens/Arize also count).
AI observability tooling: Langfuse, LangSmith, Arize Phoenix, or equivalent.
Strong AWS + CloudWatch — logs, metrics, traces, alarms.
Strong Python and/or TypeScript/JavaScript — writing real, maintainable code.
API testing (pytest, Postman, Bruno, or equivalent) and SQL for data validation.
Experience with async/event-driven systems (queues, retries, webhooks) and CI/CD integration.
Nice to Have Playwright/TypeScript for E2E, auth/permissions/tenant testing, contract testing (Pact), security/performance testing, BrowserStack, billing/financial domain experience, phased rollouts and rollback planning.
📌 AI Quality Engineer (México)
🏢 Cafeto Software
📍 México