24 sep
|
Aptonet
|
México
Job Summary
As a Senior Data Engineer, you will be responsible for designing, building, and optimizing big data pipelines. You will work closely with data scientists and other stakeholders to support the data needs of the organization.
Responsibilities
- Design and implement high-velocity, high-volume data streaming solutions using Apache Kafka and Spark Streaming.
- Develop real-time data processing and streaming techniques using Spark Structured Streaming and Kafka.
- Troubleshoot and optimize Spark applications for peak performance.
- Work closely with Python and/or Scala for data processing tasks (PySpark/Scala-Spark).
- Utilize Databricks for cloud-based big data solutions.
- Build, test, and optimize big data ingestion pipelines, architectures, and datasets.
- Deploy data platforms on Azure or AWS and manage serverless technologies like S3, Kinesis/MSK, Lambda, and Glue.
- Operate messaging platforms like Kafka, Amazon MSK, TIBCO EMS, or IBM MQ Series for asynchronous data communication.
- Manage Databricks Notebooks, work with Delta Lake using both Python and Spark SQL, and manage Delta Live Tables and the Unity Catalog.
- Ingest data from various formats including JSON, XML, and CSV.
- Work with NoSQL databases like HBASE and/or Cassandra.
- Perform shell scripting and other tasks on Unix/Linux platforms.
- Work with other database solutions like Kudu/Impala or Delta Lake.
Qualifications - must have
- Must have hands-on experience with high-velocity, high-volume stream processing: Apache Kafka and Spark Streaming.
- Experience with real-time data processing and streaming techniques using Spark structured streaming and Kafka.
- Deep knowledge of troubleshooting and tuning Spark applications.
- Must have hands-on experience with Python and/or Scala (PySpark/Scala-Spark).
- Must have experience with Databricks.
- Must have hands-on experience building, testing, and optimizing big data ingestion pipelines, architectures, and datasets.
- Experience in successfully building and deploying a new data platform on Azure/AWS.
- Experience in Azure/AWS Serverless technologies, like S3, Kinesis/MSK, Lambda, and Glue.
- Strong knowledge of Messaging Platforms like Kafka, Amazon MSK & TIBCO EMS or IBM MQ Series.
- Experience with Databricks UI, Managing Databricks Notebooks, Delta Lake with Python, Delta Lake with Spark SQL, Delta Live Tables, Unity Catalog.
- Experience with data ingestion of different file formats like JSON, XML, CSV.
- Experience with NoSQL databases, including HBASE and/or Cassandra.
- Knowledge of Unix/Linux platform and shell scripting is a must.
- Experience with database solutions like Kudu/Impala, or Delta Lake.
📌 Senior Data Engineer (México)
🏢 Aptonet
📍 México