07 ago
|
Improving
|
Xico
Data Lake & Warehouse ArchitectureA multi-tenant data lake and warehouse that unifies all research data — surveys, qualitative feedback, CRM, media transcripts, and more — in a structured, AI-consumable formatTenant-isolated data architecture where enterprise client data is structurally separated at the storage and query layerProvenance-aware data models where every data point carries full traceability back to its sourceData Ingestion & PipelinesBatch ingestion pipelines that migrate and continuously sync data from existing relational databases and cloud storage into the new lake architectureA nightly profile enrichment pipeline that rebuilds living user profiles from all data sources within each client accountData Access & AI EnablementData access layers serving AI agents via MCP, qualitative search via RAG pipelines, statistical computation tools, REST APIs, and bulk export© Get on Board.Your Success MetricsQuickly integrates with the engineering team and contributes meaningfully to the data platform buildTakes ownership of assigned pipeline and infrastructure work end-to-end, from design through productionBrings architectural recommendations and solutions proactively, rather than waiting for directionDemonstrates strong collaboration and communication across engineering and product teamsWhat you'll bring:You have 5+ years of deep, hands-on experience building production lakehouses on Databricks.
You write clean PySpark and Python, model data thoughtfully, and know how to build for a multi-tenant SaaS environment.Deep production experience across the Databricks platform including Unity Catalog, Delta Live Tables, Databricks SQL, and WorkflowsDelta Lake as a production table format — ACID transactions, schema evolution, performance optimization,
and multi-tenant governance via Unity Catalog Experience building and maintaining dbt transformation projects using the Databricks adapter in a production environmentPySpark for large-scale data transformation and batch pipeline authoringStrong understanding of batch ingestion pipeline design — migrating from relational sources like MySQL and PostgreSQL into a lakehouse architectureExperience with a modern pipeline orchestrator such as Dagster, Prefect, or Databricks Workflows; Dagster experience is a strong positiveFamiliarity with vector databases, embedding pipelines, and RAG patterns for AI workloads — using tools such as Databricks Vector Search, pgvector, or Amazon OpenSearchExposure to AI agent and LLM-serving infrastructure including Amazon Bedrock, AgentCore, and StrandsExperience with data cataloging and governance tools such as Unity Catalog or OpenMetadataData modeling for multi-tenant analytical workloads — partitioning strategy, schema design, and tenant isolation patternsDatabricks on AWS — workspace configuration, S3 integration, IAM, and cost governanceInfrastructure as code using Databricks Asset Bundles or TerraformStrong Python and SQL skillsPreferred (but not required):Databricks certifications — Data Engineer Associate or ProfessionalSalesforce or CRM data integration experiencePrior experience in a multi-tenant SaaS environment with strict data isolation requirementsExperience migrating from OLTP to a lakehouse architectureAWS-Native Experience — A Strong PositiveCandidates with experience in AWS-native data services are strongly valued.
Engineers who understand both Databricks and AWS-native approaches bring a broader architectural perspective that helps the team make better long-term platform decisions.Apache Iceberg, AWS Glue, Athena, and DynamoDB experience*** Please note that our offices are located in Guadalajara and Aguascalientes.
Candidates from other cities are welcome to join us through our remote work model.
**Awesome Benefits :BenefitsMajor Medical Expense InsuranceLife InsuranceDental and Vision InsuranceMental Health SupportIMSS (Mexican Social Security)Seniority BonusSavings Fund ProgramCareer Development PlanChristmas Bonus (Aguinaldo)Vacation BonusCorporate Retirement PlanCertifications and Training ProgramsInternal EventsTotalPass Wellness ProgramAdditional Paid Time OffAdditional Protection and DiscountsAuto and Motorcycle InsurancePet InsurancePersonal Belongings InsuranceAnd much more!!
!*** Please note that our offices are located in Guadalajara and Aguascalientes.
Candidates from other cities are welcome to join us through our remote work model.
****GETONBRD Job ID: *****Relocation offeredIf you are moving in from another country, Improving helps you with your relocation.Pet-friendlyPets are welcome at the premises.Adaptable hoursFlexible schedule and freedom for attending family needs or personal errands.Partially remoteYou can work from your home some days a week.Health coverageImproving pays or copays health insurance for employees.Computer providedImproving provides a computer for your work.Informal dress codeNo dress code is enforced.Vacation over legalImproving gives you paid vacations over the legal minimum.Beverages and snacksImproving offers beverages and snacks for free consumption.Remote work policyHybridThis job is performed partly from home and partly at the office in Ciudad de México, Guadalajara, Monterrey, Aguascalientes or Querétaro (Mexico).
📌 Databricks Data Engineer (Xico)
🏢 Improving
📍 Xico