02 ago
|
Improving
|
Xico
Data Lake & Warehouse Architecture
A multi-tenant data lake and warehouse that unifies all research data — surveys, qualitative feedback, CRM, media transcripts, and more — in a structured, AI-consumable format
Tenant-isolated data architecture where enterprise client data is structurally separated at the storage and query layer
Provenance-aware data models where every data point carries full traceability back to its source
Data Ingestion & Pipelines
Batch ingestion pipelines that migrate and continuously sync data from existing relational databases and cloud storage into the new lake architecture
A nightly profile enrichment pipeline that rebuilds living user profiles from all data sources within each client account
Data Access & AI Enablement
Data access layers serving AI agents via MCP, qualitative search via RAG pipelines, statistical computation tools, REST APIs, and bulk export
Apply to this job without intermediaries on Get on Board.
Your Success Metrics
Quickly integrates with the engineering team and contributes meaningfully to the data platform build
Takes ownership of assigned pipeline and infrastructure work end-to-end, from design through production
Brings architectural recommendations and solutions proactively, rather than waiting for direction
Demonstrates strong collaboration and communication across engineering and product teams
What you'll bring:
You have 5+ years of deep, hands-on experience building production lakehouses on Databricks. You write clean PySpark and Python, model data thoughtfully, and know how to build for a multi-tenant SaaS environment.
Deep production experience across the Databricks platform including Unity Catalog, Delta Live Tables, Databricks SQL, and Workflows
Delta Lake as a production table format — ACID transactions, schema evolution, performance optimization,
and multi-tenant governance via Unity Catalog Experience building and maintaining dbt transformation projects using the Databricks adapter in a production environment
PySpark for large-scale data transformation and batch pipeline authoring
Strong understanding of batch ingestion pipeline design — migrating from relational sources like MySQL and PostgreSQL into a lakehouse architecture
Experience with a modern pipeline orchestrator such as Dagster, Prefect, or Databricks Workflows; Dagster experience is a strong positive
Familiarity with vector databases, embedding pipelines, and RAG patterns for AI workloads — using tools such as Databricks Vector Search, pgvector, or Amazon OpenSearch
Exposure to AI agent and LLM-serving infrastructure including Amazon Bedrock, AgentCore, and Strands
Experience with data cataloging and governance tools such as Unity Catalog or OpenMetadata
Data modeling for multi-tenant analytical workloads — partitioning strategy, schema design, and tenant isolation patterns
Databricks on AWS — workspace configuration, S3 integration, IAM, and cost governance
Infrastructure as code using Databricks Asset Bundles or Terraform
Strong Python and SQL skills
Preferred (but not required):
Databricks certifications — Data Engineer Associate or Professional
Salesforce or CRM data integration experience
Prior experience in a multi-tenant SaaS environment with strict data isolation requirements
Experience migrating from OLTP to a lakehouse architecture
AWS-Native Experience — A Strong Positive
Candidates with experience in AWS-native data services are strongly valued.
Engineers who understand both Databricks and AWS-native approaches bring a broader architectural perspective that helps the team make better long-term platform decisions.
Apache Iceberg, AWS Glue, Athena, and DynamoDB experience*** Please note that our offices are located in Guadalajara and Aguascalientes. Candidates from other cities are welcome to join us through our remote work model. **
Awesome Benefits :
Benefits
Major Medical Expense Insurance
Life Insurance
Dental and Vision Insurance
Mental Health Support
IMSS (Mexican Social Security)
Seniority Bonus
Savings Fund Program
Career Development Plan
Christmas Bonus (Aguinaldo)
Vacation Bonus
Corporate Retirement Plan
Certifications and Training Programs
Internal Events
TotalPass Wellness Program
Additional Paid Time Off
Additional Protection and Discounts
Auto and Motorcycle Insurance
Pet Insurance
Personal Belongings Insurance
And much more!!!
*** Please note that our offices are located in Guadalajara and Aguascalientes. Candidates from other cities are welcome to join us through our remote work model. ****
GETONBRD Job ID: *****
Relocation offered
If you are moving in from another country, Improving helps you with your relocation.
Pet-friendly
Pets are welcome at the premises.
Versátil hours
Flexible schedule and freedom for attending family needs or personal errands.
Partially remote
You can work from your home some days a week.
Health coverage
Improving pays or copays health insurance for employees.
Computer provided
Improving provides a computer for your work.
Informal dress code
No dress code is enforced.
Vacation over legal
Improving gives you paid vacations over the legal minimum.
Beverages and snacks
Improving offers beverages and snacks for free consumption.
Remote work policy
Hybrid
This job is performed partly from home and partly at the office in Ciudad de México, Guadalajara, Monterrey, Aguascalientes or Querétaro (Mexico).
📌 Databricks Data Engineer (Xico)
🏢 Improving
📍 Xico