Principal SRE
Alpheya · Abu Dhabi, Abu Dhabi Emirate, United Arab Emirates
Apply & track with Apply EdgeAbout AlpheyaAlpheya is a wealth-management technology company headquartered in Abu Dhabi. Banks are our customers. We build each one a tailored investing experience, from mobile apps to advisor and back-office portals, on a single shared platform covering the full order-to-custody lifecycle. The platform runs as SaaS on Microsoft Azure, with on-premises delivery for banks that require it.The roleOver the next six months we are taking multiple banks live as SaaS customers. Our engineering teams build the platform and our SREs run the infrastructure. What we don't yet have is a single leader accountable for the service our customers buy: the SLAs, the incident review a bank CIO sits in on, the audits, the disaster-recovery program, the cost of running each tenant. That is this role.You will report to the CTO and own SaaS operations end to end, from the Kubernetes clusters to the quarterly service review with a bank's executives. The mandate: make onboarding the fifth bank a checklist instead of a project.What You'll OwnService management for every SaaS customer. SLA definition and reporting, incident communications and post incident reviews, security questionnaires, audit cycles, and the day-to-day support model with each bank's service deskRelease and deployment operations. The release calendar across the tenant estate, environment promotion and rollback discipline, production change management, and coordinating rollouts with each bank's change and freeze windows. Engineering builds the pipelines; you decide when and how production changesThe reliability program. Tested disaster recovery and business continuity, backup verification, patching and vulnerability-management cadence, capacity planning, and a cost-per-tenant model the CFO can useThe tenant onboarding runbook, so new banks go live on a repeatable pathOur European launch. Two new Azure regions with data residency, DR, and support coverage in place, operated rather than merely deployedCompliance operations for outsourcing arrangements, working with bank risk teams under frameworks such as ISO 27001, SOC 2, and European and Gulf outsourcing regulationYour first six monthsFour bank go-lives across two geographies, each with agreed SLAs, escalation paths, and incident procedures in place from day oneA disaster-recovery exercise run and documented for at least one production environmentAn on-call rotation covering both regions without heroicsA tenant cost model and a capacity plan for the whole estateRequirements12+ years in production operations, several of them leading the function for multi-tenant SaaSOwnership of customer-facing service management. You have led a severity-one bridge, presented the post-incident review to a customer's executives, and been through their auditsYou have built or scaled an SRE or platform-operations team and designed on-call across regionsEnough Kubernetes and cloud depth to challenge your engineers on failure modes, DR design, and capacity claimsWorking knowledge of Azure itself. Regions and availability zones, networking and private connectivity, identity, and quota planning. You will negotiate these with Microsoft and with bank security teamsExperience with certification and audit regimes: ISO 27001, SOC 2, or bank outsourcing regulation in the EU or the GulfNice to haveWealth management, brokerage, or capital-markets domain exposureExperience operating streaming or workflow infrastructure (Temporal, Kafka, or similar)Experience with on-premises software delivery to enterprise customers