AI Application Engineer (Part-Time) (Ciudad de México)

AI Application Engineer (Part-Time) (Ciudad de México)

12 ago
|
Workana
|
Ciudad de México

12 ago

Workana

Ciudad de México

Client: medxprts.ai
Location: Remote
Type: Part-time, with potential to transition to Full-time
Schedule: U.S. timezone overlap required
Description Medxprts.ai is building an AI-powered platform for the legal and healthcare space, using LLMs, agentic workflows, and automation to create production-grade applications.
We are seeking an AI Engineer - LLM Fine-Tuning: a hands-on engineer who has personally trained and fine-tuned open-weight models, built training infrastructure, engineered datasets from messy documents, and established rigorous evaluation and preference/feedback training pipelines. This role works closely with engineering teams to ship models and integrated features into production.

Responsibilities Design, implement, and run LLM fine-tuning experiments (LoRA/QLoRA and full SFT) on open-weight models (e.g., Llama, Mistral, Qwen) and ship trained models into product workflows.
Build and maintain training infrastructure using PyTorch and Hugging Face tooling (Transformers, PEFT, TRL/Axolotl), including multi-GPU training orchestration (DeepSpeed/FSDP) on AWS or GCP.
Engineer datasets from real-world unstructured sources (long PDFs, medical/legal records), performing deduplication, filtering, contamination checks, and train/eval splits.
Create evaluation harnesses tailored to domain needs: held-out test sets, LLM-as-judge with human calibration, regression tests across model versions, and automated monitoring for model drift.
Implement preference/feedback training workflows (DPO/RLHF-style or similar) to learn from expert corrections (doctor-in-the-loop),



and integrate feedback loops into model retraining pipelines.
Collaborate with backend/frontend engineers to integrate fine-tuned models into services, optimize inference latency/cost, and support production deployments.
Participate in PR reviews, release processes, incident debugging, and continuous improvement of training and deployment tooling.
Ensure secure, compliant handling of sensitive data (HIPAA-awareness is highly preferred) during dataset preparation and model training.
Requirements
Hands-on experience fine-tuning LLMs: personally trained or fine-tuned open-weight models using LoRA/QLoRA and full SFT; able to explain trade-offs and provide at least one shipped example.
Training infrastructure experience: PyTorch + Hugging Face ecosystem (Transformers, PEFT, TRL/Axolotl), multi-GPU training knowledge (DeepSpeed or FSDP), and running training workloads on AWS or GCP.
Dataset engineering expertise: built instruction/preference datasets from messy, unstructured documents; practical knowledge of deduplication, filtering, train/eval splits, and contamination prevention.
Strong evaluation discipline: designed domain-specific evaluation harnesses beyond standard benchmarks, including human-calibrated judge setups and regression testing.




Practical experience with preference/feedback learning methods (DPO, RLHF-style workflows, or equivalent) and integrating expert feedback into model updates.
Solid software engineering fundamentals: production workflows (Git, PRs), debugging, testing, deployment experience, and maintainable code.
Experience with APIs, databases, and service integration for model inference.
Ability to work independently, learn quickly, and follow technical direction.
Good English communication skills and availability to overlap with U.S. working hours.

Nice to Have Experience with long-context handling strategies for very large documents (retrieval-aware training, context extension, RAG for multi-thousand-page sources).
Model deployment and inference optimization: quantization (GPTQ/AWQ), vLLM/TGI serving, batching/throughput tuning, latency and cost optimization.
Familiarity with vector databases, advanced RAG pipelines, MCPs, or n8n-style workflow automation tools.
Knowledge of Docker, CI/CD, Kubernetes/EKS, or serverless infrastructure.
Prior experience in healthcare, legaltech, HIPAA-aware processes, or other regulated/data-sensitive environments.
Public portfolio, GitHub, or examples of shipped LLM/agentic applications and fine-tuning projects.
Benefits
Fully remote role.
Opportunity to work on real AI products in the legal and healthcare domain.
High ownership and autonomy.
Performance-based incentives and outcome-driven bonuses.
Potential to grow into a long-term, full-time role.

📌 AI Application Engineer (Part-Time) (Ciudad de México)
🏢 Workana
📍 Ciudad de México

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: ai application engineer (part-time) (ciudad de méxico) / ciudad de méxico

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: ai application engineer (part-time) (ciudad de méxico) / ciudad de méxico