• Build robust golden datasets extracted from UAT logs and utilize frontier models to synthetically generate variations (typos, phrasing, syntax) for robust testing.
• Construct local developer sandbox environments using the Google ADK framework with version-controlled prompts in Cider.
• Generate test scripts and code snippets mapped to defined success metrics for continuous local unit testing.
2. Production Pipeline Automation & System Architecture
• Architect and deploy language-agnostic RPC endpoints to systematically invoke GTM agents within the ecosystem.
• Build resilient production pipelines supporting parallel inference execution across 1,000+ trajectory datasets in under 15 minutes.
• Stand up a centralized Model Context Protocol (MCP) logging server to capture raw prompts, tool trajectories, SQL queries, and token costs using strict JSON schemas.
• Implement asynchronous message queues (Pub/Sub) for rate-limiting/backpressure,
along with retry policies for network and generation failures.
• Establish CI/CD Pull Request (PR) gates that block commits causing capability regressions, and enable pre-production shadow deployments using production mirror traffic.
3. Skill Benchmarking & Trajectory Validation
• Author eval test suites to isolate specific agent competencies (e.g., CRM writes, SQL analytics).
• Inject sandboxed mocks to validate tool-calling logic without producing live CRM side effects or executing heavy database reads.
• Validate multi-turn trajectories, checking chronological tool order, API loop prevention, and exact parameter payload assertions (e.g., date ranges, seller regions).
• Stream execution logs for baseline delta analysis and capability scoring.
• Build low-latency Hydra ETL pipelines to stream structured evaluation JSON records in
📌 Lead AI/ML Engineer (Guadalajara)
🏢 Ampstek
📍 Guadalajara
Postulate a este anuncio
Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.