09 sep
|
Caylent
|
Puerto Vallarta
09 sep
Caylent
Puerto Vallarta
Caylent is a cloud native services company that helps organizations bring the best out of their people and technology using Amazon Web Services (AWS). We provide a full-range of AWS services including workload migrations and modernization, cloud native application development, DevOps, data engineering, security and compliance, and everything in between. We are a general company and operate fully remote with employees in Canada, the United States, and Latin America.
This is a senior role for someone who leads from both directions at once — deeply technical on customer engagements, and fully accountable for the growth and performance of a team of ML engineers and architects. Run regular structured 1:1s, provide candid feedback at meaningful milestones, and actively invest in each person's growth — whether they are early in their career or highly experienced.
Manage performance: Recognize strong contributors and address performance gaps directly and early. Evaluate customer environments end-to-end — infrastructure, data pipelines, model lifecycle, and organizational readiness — and produce recommendations that drive executive decisions and open the door to the next engagement. Serve as the senior technical authority on engagements, setting architectural direction, ensuring technical quality across the team, and making the calls that matter when tradeoffs are hard. Help customers build ML systems they can actually own and sustain — translating MLOps, LLMOps, and production monitoring complexity into standards their engineering teams can execute and their leadership can act on. Drive architecture and solution design from kickoff through delivery — setting technical direction, unblocking the team on hard problems, and ensuring the work meets Caylent's quality standards.
Depending on the engagement, you are either the primary client contact owning all architect-level outcomes, or the senior technical authority providing oversight across the team. Demonstrated people management experience — hiring, performance calibration, career development, and the ability to have difficult conversations directly and constructively. Deep, current knowledge of the AWS ML and GenAI ecosystem, with the ability to make and defend architectural decisions across the full ML lifecycle — from data and feature engineering through training, deployment, and monitoring. Proven ability to architect and govern production ML systems end-to-end, translating MLOps, LLMOps, and broader AI operations complexity into standards that engineering teams can execute and executives can act on. Deep expertise across foundation model adaptation — fine-tuning (LoRA, QLoRA, PEFT), alignment (RLHF, DPO), inference optimization, and distributed training — combined with RAG and agentic system design, including multi-agent architectures, MCP integration, and human-in-the-loop patterns on AWS. AWS Certified Machine Learning – Specialty and/or AWS Certified Solutions Architect – Professional. Deep fluency in responsible AI practices — model evaluation, bias detection, fairness frameworks, and AI governance — applied in enterprise deployments.
Fluency in AIOps patterns — designing agentic workflows for anomaly detection, automated root cause analysis, and remediation across observability platforms — and the ability to translate AI operations outcomes into measurable business value for customers. Candidates are expected to prescribe — not just recognize — with the judgment to maximize what AWS makes possible and the experience to know how open-source tooling strengthens it. Classical ML, Computer Vision, NLP, Generative AI &
• LLMs, AI Agents &
• Autonomous Systems, Intelligent Document Processing, Video Understanding, Speech &
• Audio, Time Series &
• Forecasting, Recommender Systems, Graph ML, Reinforcement Learning, Multimodal AI AWS ML Platform: Bedrock, Anthropic API, OpenAI API, Google Gemini API, Azure OpenAI — with the judgment to reason across provider tradeoffs in enterprise contexts AWS AI Services: Rekognition, Comprehend, Transcribe, Textract, Translate, Personalize, Neptune, Kinesis Video Streams, Polly Data Platform: Apache Spark / PySpark, Apache Kafka, Amazon Kinesis, Apache Iceberg, Delta Lake, Apache Hudi, AWS Glue Vector Databases: PyTorch, TensorFlow, JAX, Scikit-learn, XGBoost, HuggingFace (Transformers, PEFT, TRL), LangChain, LlamaIndex, DSPy, Ollama MLflow, W&B;, Airflow / MWAA (data orchestration), Dagster (asset-based pipelines), Kubeflow Pipelines, CI/CD, IaC (CloudFormation, CDK, Terraform), Docker, Kubernetes, ML Governance (lineage, data contracts, audit), Responsible AI / Bias &
• Fairness 100% remote work Medical Insurance for you and eligible dependents Generous holidays and flexible PTO Equipment &
• Office Stipend Individual professional development plan NOTE: We're unable to provide visa sponsorship now or at any time in the future.
📌 Senior Manager, DevOps Engineering (Puerto Vallarta)
🏢 Caylent
📍 Puerto Vallarta