Senior Data Engineer
ALPHA10X · Dubai, Dubai, United Arab Emirates
Apply & track with Apply EdgeSenior Data EngineerDoha, QatarAbout this role:ALPHA10X is an award-winning, AI-native Decision Intelligence platform built for private markets - a market projected to exceed $65 trillion in AUM by 2032, but one in which investment decisions remain constrained by fragmented data, subjective judgment and limited predictive infrastructure. At its core is ALPHAcodex, a governed reasoning system that combines data, inference, institutional memory and continuous learning to generate quantitative signals and probabilistic foresight. The platform is designed to improve underwriting and capital allocation over time: each decision becomes institutional memory, strengthening the intelligence applied to subsequent decisions and creating the potential for a compounding advantage in judgment, conviction and investment performance.Founded in France, ALPHA10X has progressively shifted its center of gravity to the Middle East as the region has become increasingly important to global capital allocation. The company is establishing its global headquarters and principal operating hub in Doha, providing a strategic base for product development, commercial expansion and global growth.ALPHA10X has raised $25 million in seed capital from prominent investors and academics and completed more than 70 engagements across 11 countries. The company is now entering its next phase of scale, strengthening leadership, accelerating R&D, building its Doha organization and expanding institutional market penetration ahead of a planned Series A financing.The Senior Data Engineer owns the data side of our client engagement in Doha. Joining our founding team in Doha and working day to day with our Senior Solutions Architect, you will bring the client’s own data and new sector-specific sources into the ALPHA10X Ledger, extend our data model (ontology) to cover them, and make that data reliable enough for our AI workflows to reason on.You will build and run pipelines in Databricks and Azure Data Factory on our Azure-hosted platform, and uphold agile engineering practices: unit testing, automation, and continuous integration via CI/CD.The successful candidate will be someone who takes initiative, stays curious and has an unrelenting drive to push the boundaries of AI and data science. They will be adaptive and introspective: willing to learn, guide, lead and follow, with the low ego of someone for whom the outcome matters more than who gets the credit. This is an opportunity to play a key role in advancing the capabilities of AI in fintech, working with some of the brightest minds in the industry, as one of the first engineers in our Doha office, shaping how ALPHA10X delivers for clients in the region. As part of the future-shaping team, you will have a direct impact on the company’s growth and strategic direction, in a fast-paced, high-growth startup with significant equity participation and potential for personal and professional advancement, as part of a team that connects people, capital and ideas to solve some of the world’s greatest challenges through AI-driven financial solutions.
Key Responsibilities
- ETL Development: Design, develop and implement robust, efficient and scalable data pipelines into the ALPHA10X Ledger. Ingest structured feeds, APIs and unstructured documents. Build batch pipelines for the data we pre-populate. Build on-demand enrichment pipelines triggered when a user asks for something. Build and run pipelines in Databricks and Azure Data Factory, with unit testing, automation and CI/CD. The expectation is that every pipeline is production-grade from the start: tested, automated and repeatable.· Client Data Integration: Onboard the client’s proprietary and licensed datasets into our platform. Put data contracts, schema mapping, secure connectors and permissioned access in place. Work within the data-residency and security requirements agreed with the client. Client data is brought in on the client’s terms: secure, permissioned and within agreed residency rules. · Data Modelling & Ontology: Extend our ontology to cover each new source. Map each new source onto our ontology (organizations, people, transactions, funds, and the relationships between them). Decide with the Head of Data whether a new concept should be an entity or an attribute. Register new entities, relationships and attributes in our field catalogue so they are exposed through the Ledger API. Build the data layer for sector-specific ontology packs: new source types (e.g. patents, scientific literature, regulatory and project registries) and new entity types, driven by the client’s priority use cases. New data only creates value once it is modelled consistently and exposed through the Ledger API. · Entity Resolution: Extend our record-linkage capability to new sources. Ensure client and third-party records resolve to the right organization. Measure matching precision and recall rather than assuming them. Match quality is measured, not assumed. · Data Quality & Observability: Define and enforce quality checks on every pipeline. Check completeness, freshness, duplicates and schema conformance. Put alerting and dashboards in place. Issues are caught before they reach a client-facing workflow. · Cost & Scale Trade-offs: Make every new source a deliberate economic decision. Estimate the cost to acquire, process and refresh each new source. Recommend whether it should be pre-populated in batch or retrieved on demand. No source is added without a clear view of what it costs to keep it current. · Documentation & Handover: Make what you build operable by others. Document pipelines, data contracts and runbooks. Ensure they can be operated by the data team in France and by future Doha hires. The data layer should run without depending on any one person. What Success Looks Like Success in this role will be visible in the reliability of the data the client’s workflows run on. · Within the first 30 days, the Senior Data Engineer will have:o been onboarded on our platform, Ledger and pipelines; ando mapped the client’s data sources and assessed their readiness (what exists, what is licensed, what is missing).· Within the first 60 days, the Senior Data Engineer will have:o ingested sample client data and assessed its quality; ando drafted an initial ontology for the client’s domain and reviewed it with their experts. · Within the first 6 months, the Senior Data Engineer will have:o production pipelines running on the client’s agreed sources;o entity resolution, monitoring and data isolation in place; ando data operations running steadily as the platform goes live. What We Are Looking ForThis role requires a combination of hands-on data engineering at scale, rigour in data quality and comfort working directly with a client’s data. We are looking for someone who:· Has a Master’s degree (minimum) in computer science or engineering.· Has 5+ years of significant experience (internships excluded) in Data Engineering (not Data Science, nor Data Analysis).· Has excellent verbal and written English communication skills.· Has outstanding programming skills in PySpark, SQL and Python.· Has proven experience processing data at terabyte scale in production.· Is comfortable working with Git.· Is familiar with document databases (CosmosDB, MongoDB), graph databases (Neo4J) and ELK.· Has hands-on Azure Cloud services experience (Databricks, Azure Data Factory and Azure DevOps).· Has a strong understanding of SQL and NoSQL database internals, query languages and distributed computing fundamentals.· Has built or run entity resolution / record linkage across messy sources (fuzzy matching, blocking, measuring match quality).· Has designed schemas or ontologies for a knowledge graph or linked data, including entity–relationship modelling.· Has delivered data integration for an external client or partner, including data contracts, secure transfer and access controls.· Learns fast and has a continuous-improvement attitude.· Is team-work oriented.· Is proactive: raises data issues and trade-offs early, with a recommendation.· Is comfortable working directly with client teams: explaining data issues to non-technical stakeholders and asking the right questions about their data.· Can work autonomously in a small on-site team, in close contact with the wider data team in France.· Experience in regulated or sovereign environments with data-residency requirements, or with financial, private-markets, or energy and sustainability data, would be valuable. However, a track record of reliable data engineering in production matters more than any particular domain.Our Tech StackWhat you will work with day to day:· Python· LangGraph, LangChain / LangSmith· ElasticSearch· Azure Cloud· Claude / Anthropic models and managed agentsThe wider platform your work plugs into:· TypeScript, Angular / NestJS, Cypress