Senior Data Engineer
DigyCorp · Thuwal, Makkah, Saudi Arabia
Apply & track with Apply EdgeRole OverviewThis is an amazing opportunity for the right person. You will lead the data engineering for a petabyte scale industrial digital twin, at the center of a national initiative delivering the world's largest coral restoration programme, working with imagery, geospatial data and sensor telemetry that few data engineers see in one place. Beyond conventional pipeline work you will build the curated datasets that feed production AI and MLOps pipelines, and own the data governance framework from first principles, in a small senior team where your decisions reach production quickly.This is a hands-on delivery-first role combined with advisory responsibility pertaining to data. You will work with the Databricks Lakehouse (bronze/silver/gold), build and run Azure Data Factory orchestration, manage the real-time ingestion chain from Event Hubs through Function Apps and Batch Accounts, and migrate operational data from source systems. You will also own the continued functional enhancements and optimization of the live application data architecture. This can include schema changes as well as performance tuning.Alongside deep hands-on engineering, you will be the person the team and the client look to for direction on data architecture and data governance: shaping the digital twin platform medallion and application architecture, and standing up and operating the governance framework the engagement commits to the data domain map, ownership register, data quality rules and issue routing, classification-driven access, metadata and lifecycle.Key ResponsibilitiesDatabricks Lakehouse & Pipelines
- Build and maintain bronze/silver/gold pipelines in Databricks, processing CSV and other source formats into curated, analysis-ready datasets.
- Develop Databricks notebooks/jobs (PySpark/SQL), manage job scheduling for recurring and incremental loads, and tune performance.
- Register curated datasets in Unity Catalog with correct metadata, ownership and access tags, per the architecture team's cataloguing standards.
- Apply data modeling and data warehousing best practices (dimensional modeling, star schema, slowly changing dimensions) when designing silver/gold layer schemas.
- Use Databricks Lakebase — the managed, Postgres-compatible operational database built on the Lakehouse — where a workload needs transactional/OLTP-style access alongside the Delta Lake data.Orchestration & Ingestion
- Build and maintain Azure Data Factory pipelines and triggers that move data from source systems into Databricks — this movement is trigger-driven, not automatic.
- Consume streaming data from Azure Event Hubs across multiple concurrent sources, and configure triggers so Databricks pulls new/changed data as it arrives.
- Develop Azure Function Apps that fire on incoming Event Hub data and initiate downstream processing via Batch Accounts.
- Manage Azure Batch Accounts that receive data from third-party tools, IoT/sensor feeds and other external sources, and move it into Databricks once triggered.
- Keep the ingestion chain (Event Hub → Function App → Batch Account → Databricks) resilient, monitored, and recoverable from failure.Migration, Security & Storage
- Migrate data from the different sources/platforms into Databricks, preserving data integrity and minimizing disruption to source systems.
- Support migration of large imagery/scientific-archive data into the platform, including metadata preservation and integrity validation.
- Manage credentials and secrets in Azure Key Vault, and implement access controls (RBAC, managed identities) per architecture guidance.
- Maintain Azure Data Lake as the platform's storage layer, including file and metadata organisation across Databricks and Batch Account outputs.Downstream Enablement
- Prepare curated, governed datasets that feed Power BI/Tableau reporting and Digital Twin applications.
- Prepare feature-ready datasets for AI/ML consumption, and support MLflow-tracked training and inference data needs.
- Build and maintain the feature store and feature-engineering pipelines that support the programme's ML cadence, working with the ML Lead so models consume governed, versioned datasets rather than ad-hoc extracts.
- Provide the data foundation for MLOps — reproducible training and inference datasets, dataset and feature versioning, lineage back to the gold layer, and drift/freshness monitoring that supports model retraining.
- Build the data pipelines behind GenAI and RAG use cases: preparing and chunking documents and scientific/operational text, generating and maintaining embeddings, managing vector indexes, and keeping retrieval sources refreshed and access-controlled so retrieval respects the same classification rules as the rest of the platform.
- Enable advanced analytics capabilities on the platform, including graph and knowledge-graph data structures that support simulation, prediction and decision support alongside the Digital Twin.
- Work with geospatial (PostGIS) and time-series/sensor data feeds as part of the ingestion and curation pipeline.Data Architecture
- Own the Lakehouse data architecture in practice — medallion layering, domain boundaries, curated data products and serving patterns — documented in the programme's digital blueprint, and set the integration and data standards any contributor, internal or third-party, must meet before data enters the governed layers.
- Own application level data architecture across all current and future applications.
- Advise on the open target-state decisions and work with the Platform Lead and ML Lead so ingestion, curation, reporting and ML consumption are architected as one estateData Governance Leadership
- Run the initial data assessment and define the governance framework from it — the data domain map, the ownership register with a named owner per domain, and the standards covering data quality, classification-driven access, privacy, metadata and lifecycle.
- Implement that framework in the platform rather than on paper: data quality rules with live issue routing to the accountable domain owner, the QA/QC split with the business, and retention, lifecycle and deletion processes operating for governed data.
- Own the implementation of Data Quality rules onto relevant applications. Direct application development team on what data quality controls need to be implemented in application and any future applications.
- Bring each newly onboarded domain under the framework and the metadata standard as it lands, evidence that governance is actually operating,Core Technology Skills
- Azure Databricks, Apache Spark/PySpark, Delta Lake, Unity Catalog, Medallion (bronze/silver/gold) architecture.
- Databricks Lakebase (managed, serverless Postgres-compatible OLTP database sharing the Lakehouse's storage layer) for operational/transactional workloads.
- Data modeling and data warehousing fundamentals — dimensional modeling, star schema, slowly changing dimensions — applied to curated Lakehouse layers.
- Azure Data Factory, ADLS Gen2, pipeline orchestration, and batch + streaming integration patterns.
- Azure Event Hubs, Azure Function Apps and Azure Batch Accounts for event-driven and trigger-based ingestion.
- Advanced SQL and strong Python/PySpark for data transformation and pipeline development.
- PostgreSQL, including PostGIS for geospatial data and exposure to time-series stores (e.g. TimescaleDB) for sensor/IoT data.
- Azure Key Vault, managed identities and RBAC for secure credential and access management.
- Data quality and observability practices: validation, profiling, freshness/completeness checks, pipeline monitoring.
- Experience preparing data for Power BI/Tableau consumption and for AI/ML pipelines (feature tables, MLflow-tracked datasets).
- Git-based development and CI/CD for data pipelines (Azure DevOps or equivalent).
- Working familiarity with geospatial/GIS data (ArcGIS-adjacent) and IoT/telemetry data formats.
- Data governance frameworks in practice — ownership and stewardship models, data quality rule design, classification-driven access, retention and lifecycle, metadata standards and catalogue management (Unity Catalog, Purview or equivalent).
- Lakehouse and data-product architecture — medallion design, domain boundaries, serving patterns, and the ability to document and defend architecture decisions to a client audience.
- Data lineage, cataloguing and observability tooling, and the ability to evidence that controls are operating rather than merely defined.
- Cost and performance optimisation of Azure data workloads (reserved compute/storage, cluster sizing, endpoint and logging optimisation).
- AI/ML data engineering — feature stores, feature engineering at scale, training/inference dataset preparation, MLflow-tracked datasets, and the data side of MLOps (versioning, lineage, drift and freshness monitoring).
- GenAI and RAG data patterns — document and text preparation and chunking, embedding generation, vector stores/indexes (Databricks Vector Search, pgvector or equivalent), retrieval evaluation, and applying data classification and access control to retrieval sources.