09 sep
|
Programming.Com
|
Azcapotzalco
09 sep
Programming.Com
Azcapotzalco
role: data engineer
location: remote (requires occasional travel)
responsibilities:
- build and operate robust data pipelines for ingestion, cleaning, and transformation using databricks, airflow, or dagster.
- develop efficient etl/elt workflows in python and sql to support both batch and streaming workloads.
- collaborate with ml and ai teams to deliver high-quality datasets for training, evaluation, and production features.
- model and maintain structured data assets (delta, parquet, iceberg) for reliability, versioning, and lineage tracking.
- implement orchestration and monitoring — schedule jobs, track dependencies, and automate recovery from failures.
- ensure data quality and compliance through validation frameworks, schema enforcement, and audit logging.
- contribute to data platform evolution — evaluate tools, standardize best practices, and improve developer experience.
- support performance and cost optimization across compute, storage, and orchestration systems.
qualifications:
- 3–6 years of experience as a data engineer or etl developer in a production environment.
- proficiency in python and sql; strong familiarity with databricks, spark,
or equivalent big-data frameworks.
- experience with workflow orchestration tools such as airflow, dagster, luigi or prefect.
- deep understanding of data modeling, data warehousing, and distributed data processing.
- knowledge of modern data lakehouse architectures (delta, parquet, iceberg).
- familiarity with ci/cd, github actions, and data pipeline testing frameworks.
- comfort working in a cross-functional environment with ml, product, and analytics teams.
nice to have:
- experience with sports, telemetry, or sensor data pipelines.
- familiarity with streaming frameworks (kafka, spark structured streaming, flink).
- general knowledge of american football, the nfl, and college football
- background in data governance, lineage, and observability tools (monte carlo, great expectations, unity catalog, openlineage).
- experience with cloud infrastructure (aws, gcp, or azure) and containerization (docker, kubernetes).
- exposure to best practices in machine-learning model management and mlops
📌 Data engineer (Azcapotzalco)
🏢 Programming.Com
📍 Azcapotzalco