19 ago
|
KMS Technology
|
Guadalajara
19 ago
KMS Technology
Guadalajara
At KMS Technology, we are dedicated to delivering cutting-edge solutions and services that empower businesses to achieve their goals. Our team is composed of highly skilled professionals who are passionate about technology and innovation. We provide a dynamic and collaborative work environment where you can grow your career and make a significant impact.
**Responsibilities**:
You will work with large, heterogeneous datasets, contribute to architecture and ingestion strategies, and ensure that Spark pipelines operate reliably, efficiently, and cost-effectively.
**Responsibilities**:
- Design, implement, and optimize **high-volume data ingestion pipelines** using **Apache Spark**, integrating internal and external data sources through both standard and custom connectors.
- Lead the development of **data integration strategies** for moving, transforming, and loading large-scale, diverse datasets using** Spark (PySpark)**across cloud environments such as **AWS, Azure, or GCP.**
- Translate complex business and technical requirements into **scalable PySpark notebooks and jobs** with clear, maintainable structure.
- Port existing ETL/ELT logic and transformation processes from legacy systems into optimized **PySpark-based implementations.**
- Implement testing, monitoring,
and **robust error-handling mechanisms** within Spark pipelines to ensure data integrity and operational reliability.
**Qualifications**:
- 5+ years of professional software development experience, with 3+ years focused on large-scale data engineering.
- Expert-level proficiency in Apache Spark (PySpark), including Spark RDDs, DataFrames, Spark SQL, and deep understanding of Spark internals (e.g., execution plans, DAGs, memory management).
- Proven experience designing and implementing large-scale, high-throughput Spark-based ingestion pipelines, using both custom and standard data integration patterns.
- Strong hands-on experience with ETL/ELT pipelines and data warehousing concepts.
- Advanced Python programming skills, particularly in data processing with PySpark and Pandas.
- Practical experience identifying and resolving Spark performance bottlenecks in distributed computing environments.
- Familiarity with Spark-based services such as Databricks, AWS EMR, Azure Synapse Analytics, or GCP Dataproc.
- Experience with Microsoft Fabric Spark workloads.
- Strong knowledge of SQL and experience with relational and NoSQL databases.
**Benefits and Perks**:
**Location**:Guadalajara,**Jalisco, Mexico (working from home - office won't be mandatory all the time, rather it will required from time to time).
📌 Senior Software Developer (Guadalajara)
🏢 KMS Technology
📍 Guadalajara