We are seeking a
Lead Data Quality Engineer
to drive rigorous data validation, SQL-based testing, and quality automation across cloud data environments. You will verify complex transformations, ensure consistency across multiple systems, and support migration and deployment work within modern data ecosystems. Join a distributed team focused on trustworthy datasets and apply today.
Execute QA validation for bulk data products within the data exchange ecosystem Validate match and append processes and ensure correct deployment into data pipelines Provide QA support for data platform migrations, ensuring data integrity and functional correctness Verify data products and data fulfillment processes for accuracy and completeness Perform data validation and comparisons across systems, including source input files to cloud data warehouse tables, table-to-table checks, and confirmation of data mappings and transformations against specifications Use and enhance quality check frameworks built with PySpark scripts to automate data validation Investigate defects through data analysis, identifying root causes in data pipelines or transformation logic Collaborate with engineering and data teams to triage issues, validate fixes,
and confirm production readiness Contribute to test automation for data validation and testing to increase efficiency and coverage Communicate findings, risks, and test results clearly to stakeholders 5+ years of experience in Data Quality Engineering Expertise in SQL, including complex joins across multiple tables and large datasets Hands-on experience with cloud data warehouses such as BigQuery, Redshift, or Synapse Analytics Working knowledge of cloud platform services across AWS, Azure, or GCP Understanding of data validation practices and automated data comparison techniques Background in data transformation, validation, and mapping verification using specifications Capability to understand and work with PySpark-based quality frameworks Strong data analysis and debugging skills, with an ability to identify defects in data processing pipelines Excellent written and verbal communication skills Upper-Intermediate English proficiency (B2) Proficiency in Python or PySpark development Background in data engineering or data pipeline testing environments
📌 Lead Data Quality Engineer (México)
🏢 Epam Systems
📍 México
Postulate a este anuncio
Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.