09 ago
|
Cafeto Software
|
México
09 ago
Cafeto Software
México
We're looking for someone to own the evaluation strategy for AI-powered workflows that handle real billing and collections outcomes — agent behavior, output quality, accuracy, and business-rule compliance.
What You'll Do
- Build automated AI evaluations (DeepEval, Ragas, Promptfoo, LangSmith, or similar) combining deterministic checks, LLM-as-a-judge, and production evidence.
- Create and maintain evaluation datasets: golden examples, regression suites, edge cases.
- Integrate evaluations into CI/CD so AI regressions get caught before release.
- Test the full workflow — APIs, backend services, async processing, third-party integrations — not just the model.
- Investigate failures using Langfuse, CloudWatch, logs, and traces to separate product defects from model behavior.
- Build targeted Playwright tests where browser-level coverage protects a critical workflow.
- Give clear release-readiness input based on evidence, not gut feel.
What You Need
- 5+ years hands-on QE, SDET, or AI Quality experience.
- Real experience testing LLM/AI/RAG/agentic systems — hallucinations, partial correctness, non-determinism.
- Hands-on building AI evaluations (DeepEval strongly preferred; Ragas/Promptfoo/LangSmith/TruLens/Arize also count).
- AI observability tooling: Langfuse, LangSmith, Arize Phoenix, or equivalent.
- Strong AWS + CloudWatch — logs, metrics, traces, alarms.
- Strong Python and/or TypeScript/JavaScript — writing real, maintainable code.
- API testing (pytest, Postman, Bruno, or equivalent) and SQL for data validation.
- Experience with async/event-driven systems (queues, retries, webhooks) and CI/CD integration.
Nice to Have
Playwright/TypeScript for E2E, auth/permissions/tenant testing, contract testing (Pact), security/performance testing, BrowserStack, billing/financial domain experience, phased rollouts and rollback planning.
📌 AI Quality Engineer (México)
🏢 Cafeto Software
📍 México