17 sep
|
Capmation
|
Xico
The Role We are seeking a Principal Data Engineer to join our Engineering Team and help build a real-time data platform that powers self-service analytics, data agents, and business-built applications.
The ideal candidate is a senior technical leader who can architect a lakehouse from the ground up, land data from a large and diverse set of source systems into a medallion (Bronze → Gold) architecture, and design real-time ingestion patterns that meet evolving business needs.
This role also shapes workspace and governance strategy in a way that supports Data Agent enablement, and partners with the Application Development team to expose curated data.
This position requires strong collaboration and leadership skills.
The Principal Data Engineer must work effectively with engineers, analysts, app developers, AI/agent teams, product stakeholders, and clients, while helping shape ways of working, mentoring data engineers, and driving high engineering quality.
The idóneo candidate should be proactive, pragmatic, and able to balance hands-on implementation with strategic decisions about platform design, modeling, governance, and cost.
Mission of the Role Real-time data: stand up streaming and near-real-time ingestion patterns where they meaningfully change the business outcome.
Self-service: enable Data Agents, business-led report and dashboard creation, and API access for business-built applications.
Trustworthy data products: deliver curated, governed Bronze, Silver and Gold layers from identified systems.
Key Responsibilities SQL Server Builds: architect and build Power BI dashboards using SQL Server and server replication Fabric Data Lake Build: architect and implement a Microsoft Fabric Data Lake using Lakehouse, OneLake, and medallion architecture, with a clear path from raw to curated layers.
Source-to-Silver Ingestion: land data from source systems (databases, SaaS, files, events, APIs)
into Bronze and Silver layers using Fabric Data Pipelines, Dataflows Gen2, notebooks, and shortcuts as appropriate to each source.
Real-Time Patterns: design and implement streaming and near-real-time ingestion using Fabric Real-Time Intelligence (Eventstream, Eventhouse/KQL) and/or Azure Event Hubs and CDC, balancing real-time goals with cost, complexity, and actual business need.
Workspace Strategy for Data Agents: define a Fabric workspace, capacity, and governance strategy that enables Data Agents (data-aware AI agents) to operate safely against curated data, including domain organization, item ownership, security, and lineage.
API Store with App Dev: partner with the Application Development team to design and operate an API store using Fabric API, exposing curated datasets to business-built applications with appropriate authentication, throttling, and contract design.
Modeling treat capacity, storage, and movement cost as first-class engineering metrics.
DataOps escalate platform-level risks early and recommend pragmatic alternatives.
Collaboration: drive alignment across data engineering, analytics, AI/agent, app dev, and business teams, acting as a unifying technical leader who resolves cross-team friction.
Curiosity: stay current on Fabric, Databricks, real-time platforms, data agents, and emerging lakehouse capabilities; use that understanding to anticipate challenges and guide innovation.
Required Qualifications Experience: 5+ years of experience in Data Engineering,
including significant time spent designing and operating production data platforms and pipelines.
Tech Stack Platform (must have one): Microsoft Fabric — Lakehouse, Warehouse, Data Pipelines, Dataflows Gen2, OneLake, shortcuts, capacities Databricks — Workspaces, Unity Catalog, Delta Lake, Workflows, Auto Loader, DLT, SQL Warehouses, Model Serving Real-Time / Streaming: Fabric Real-Time Intelligence (Eventstream, Eventhouse / KQL), and/or Databricks Structured Streaming, and/or Azure Event Hubs, Kafka, CDC (Debezium and similar) Shared: Power BI (semantic models, DAX; Direct Lake on Fabric or Power BI on Databricks SQL) PySpark SQL — T-SQL (Fabric track) and/or Spark SQL / ANSI SQL (Databricks track) Python Azure DevOps (CI/CD, YAML pipelines) Data API tooling: Fabric API, Azure API Management, or Databricks SQL endpoints / Model Serving Must have: Lakehouse Platform Required: deep, hands-on production experience in Microsoft Fabric OR Databricks — must have built a production lakehouse on one of these, not just experimented, with a strong understanding of the trade-offs between the two.
SQL Server Experience: hands-on experience with SQL Server and Server Replication with outputting data to Power Bi dashboards.
Medallion Architecture at Scale: demonstrated experience landing data from many heterogeneous source systems (databases, SaaS, files, events, APIs) into Bronze and Silver layers with strong reliability, idempotency, and lineage.
Real-Time / Streaming Ingestion: production experience with at least one streaming stack on the candidate's primary platform (Fabric Real-Time Intelligence / Eventstream / KQL, Databricks Structured Streaming, or Azure Event Hubs / Kafka with CDC), with the judgment to choose batch vs. micro-batch vs. streaming per use case.
Workspace
📌 Principal Data Engineer (Xico)
🏢 Capmation
📍 Xico