We are looking for a Senior Data Engineer to build and evolve the data platform powering our general workforce management ecosystem.
You will design, implement, and maintain scalable data pipelines that consolidate data from multiple operational systems, transform it into trusted analytical datasets, and make it available for reporting, product analytics, and business intelligence.You should be comfortable working with modern cloud-native data architectures on AWS, building reliable ETL/ELT pipelines, and designing data models optimized for analytical workloads.
This role requires a strong engineering mindset, balancing performance, scalability, data quality, and operational excellence while collaborating closely with software engineers, product teams, analysts, and data scientists.What You Will OwnDesign, build, and maintain scalable batch and streaming data pipelines using AWS-native services and distributed processing frameworksDevelop ETL/ELT workflows to ingest, consolidate, sanitize, enrich, and transform data from multiple internal and external systemsBuild and optimize AWS Data Lake solutions using Amazon S3, AWS Glue, Amazon Redshift, and Amazon Kinesis FirehoseDesign and implement distributed data processing jobs using Apache Spark, AWS Glue, Databricks, or equivalent technologiesDevelop orchestration workflows using Apache Airflow (MWAA), AWS Step Functions, or similar workflow orchestration platformsDesign analytical data models including star schemas, snowflake schemas, dimensional models, and optimized reporting datasetsOptimize Redshift performance through distribution strategies, sort keys, partitioning, workload tuning, and query optimizationBuild resilient pipelines supporting retries, idempotency, checkpointing, incremental processing, and partial failure recoveryImplement automated data quality validation, schema evolution, lineage tracking,
and governance controlsDevelop infrastructure and deployment automation using Infrastructure as Code and CI/CD pipelinesMonitor, troubleshoot, and continuously improve the reliability, scalability, and performance of the data platformCollaborate with analysts, software engineers, data scientists, and product managers to translate business requirements into scalable data solutionsParticipate in architecture discussions and contribute technical documentation, standards, and best practicesWhat We Are Looking For5+ years of professional experience building production data pipelines and cloud-based data platformsStrong experience with AWS data services including Amazon Redshift, AWS Glue, Amazon S3, and Amazon Kinesis FirehoseStrong Python programming skills for ETL development, automation, event processing, and scriptingAdvanced SQL expertise including query optimization, window functions, analytical queries, versioned migrations, rollback strategies, and warehouse tuningExperience designing scalable ETL/ELT pipelines for both batch and streaming workloadsExperience with distributed compute and storage using Apache Spark, AWS Glue, Databricks, or similar distributed processing frameworksStrong understanding of data warehousing concepts including dimensional modeling, star schemas, snowflake schemas, partitioning strategies, and analytical data structuresExperience designing end-to-end data architectures including ingestion, transformation, orchestration, and consumption layersExperience implementing workflow orchestration using Apache Airflow (MWAA), AWS Step Functions,
or equivalent orchestration toolsUnderstanding of data governance, metadata management, security best practices, IAM, encryption, and regulatory compliance considerationsExperience with Git-based collaborative development workflows, CI/CD pipelines, Infrastructure as Code, deployment approvals, versioned migrations, and safe rollback strategiesExperience monitoring and maintaining production data infrastructure, ensuring high availability, observability, data quality, and operational reliabilityStrong communication skills with the ability to explain technical concepts to business stakeholders and collaborate effectively across engineering, analytics, and product teamsNice to HaveExperience with Apache Iceberg, Delta Lake, Apache Hudi, or modern open table formatsExperience with dbt or SQL-based transformation frameworksFamiliarity with Kafka, Amazon MSK, or other streaming platformsExperience with Lakehouse architectures and modern analytical data platformsKnowledge of Terraform or AWS CloudFormationExperience with containerized data workloads using Docker and ECS/EKSExperience implementing DataOps practices and automated testing for data pipelinesFamiliarity with BI platforms such as Tableau, Power BI, Looker, or QuickSightExperience implementing data catalogs, lineage, and governance solutionsExposure to machine learning feature pipelines or data science infrastructureTech StackLayerTechnologyProgrammingPython, SQL, PySparkData ProcessingApache Spark, AWS Glue, DatabricksData StorageAmazon S3, Amazon Redshift, ParquetStreamingAmazon Kinesis Firehose, EventBridgeOrchestrationApache Airflow (MWAA), AWS Step FunctionsData ModelingStar Schema, Snowflake Schema, Dimensional ModelingInfrastructureAWS, IAM, CloudWatchIaC/CIGit, GitHub Actions, Terraform, CloudFormationObservabilityCloudWatch, Datadog (or equivalent observability platforms)GovernanceData Catalog, Metadata Management, Data Lineage
📌 Senior Data Engineer (Platform) (Xico)
🏢 Globalli
📍 Xico