02 ago
|
Ampstek
|
Guerrero
1. Foundations & Local Sandbox Development
- Build robust golden datasets extracted from UAT logs and utilize frontier models to synthetically generate variations (typos, phrasing, syntax) for robust testing.
- Construct local developer sandbox environments using the Google ADK framework with version-controlled prompts in Cider.
- Generate test scripts and code snippets mapped to defined success metrics for continuous local unit testing.
- Architect and deploy language-agnostic RPC endpoints to systematically invoke GTM agents within the ecosystem.
- Build resilient production pipelines supporting parallel inference execution across 1,000+ trajectory datasets in under 15 minutes.
- Stand up a centralized Model Context Protocol (MCP) logging server to capture raw prompts, tool trajectories, SQL queries, and token costs using strict JSON schemas.
- Implement asynchronous message queues (Pub/Sub) for rate-limiting/backpressure, along with retry policies for network and generation failures.
- Establish CI/CD Pull Request (PR) gates that block commits causing capability regressions, and enable pre-production shadow deployments using production mirror traffic.
3.
Skill
Benchmarking & Trajectory Validation
- Author eval test suites to isolate specific agent competencies (e.g., CRM writes, SQL analytics).
- Inject sandboxed mocks to validate tool-calling logic without producing live CRM side effects or executing heavy database reads.
- Validate multi-turn trajectories, checking chronological tool order, API loop prevention,
and exact parameter payload assertions (e.g., date ranges, seller regions).
- Stream execution logs for baseline delta analysis and capability scoring.
4.
Enterprise
Analytics, Governance & Security (Phase 4)
- Build low-latency Hydra ETL pipelines to stream structured evaluation JSON records into data warehouses and construct Plx analytics dashboards.
- Enforce automated PII masking/redaction layers for sensitive seller and financial data, while configuring retention and purging policies (e.g., 90-day trajectory logs).
Required Qualifications & Technical Skills
- Core Language: Advanced proficiency in Python.
- Software & Systems Architecture: Deep expertise in enterprise software architectures, distributed computing, async task processing, and load balancing.
- AI/LLM Telemetry & Evaluations: Proven experience in LLM performance telemetry, prompt engineering, LLM-as-a-Judge systems, and statistical inter-rater agreement (IRR/Kappa).
- Experience with Google ADK (or similar agent developer kits).
- Model Context Protocol (MCP) logging architectures.
- RPC service design and integration.
- CI/CD & Testing: Hands‑on experience integrating evaluation pipelines into CI/CD workflows and managing sandboxed unit testing environments.
- Data Engineering: Proficiency in building ETL pipelines (Hydra equivalent), structuring JSON telemetry schemas, and creating analytics dashboards (Plx/BI tools).
- Data Privacy: Knowledge of automated PII/data‑masking strategies and data retention protocols.
📌 Lead AI/ML Engineer (Guerrero)
🏢 Ampstek
📍 Guerrero