Senior Lead AI Architect (On-Premise, GenAI & OpenShift Stack)
Petrus Technologies KSA · Riyadh, Saudi Arabia
Apply & track with Apply EdgeRole Overview The Senior Lead AI Architect designs and owns the end-to-end blueprint for our enterprise Generative AI and infrastructure platform, specifically tailored for highly secure, air-gapped on-premise banking deployments. You will bridge the gap between cutting-edge LLMs and enterprise infrastructure, ensuring that NVIDIA accelerated compute, DataRobot, and Graph databases operate reliably on Red Hat OpenShift within uncompromising compliance, latency, and data-sovereignty boundaries.Key ResponsibilitiesAccelerated Compute Architecture: Blueprint scalable on-premise compute nodes leveraging high-density hardware (NVIDIA DGX systems, H100/B200 GPUs) and bare-metal configurations optimized for massive LLM workloads.GenAI & LLM Platform Design: Architect the integration framework for deploying foundational models natively using NVIDIA NIM (NVIDIA Inference Microservices), ensuring scalable, low-latency orchestration within secure networks.Platform Orchestration: Design the foundational architecture to host DataRobot Enterprise, enterprise Graph Databases, and microservices natively within an on-premise Red Hat OpenShift cluster environment using the NVIDIA GPU Operator.Data & Knowledge Graph Topology: Architect secure, high-throughput on-premise data pipelines that link relational core banking systems with Graph databases to fuel advanced Retrieval-Augmented Generation (RAG) pipelines.Security & Compliance: Embed rigorous encryption (at rest/in transit), access controls (RBAC/ABAC), and strict network isolation policies to meet global banking regulations (e.g., PCI-DSS, GDPR).Required Experience & Technical SkillsExperience: 8+ years in enterprise infrastructure architecture, with at least 3+ years dedicated to designing large-scale AI/ML, LLM, or GPU accelerated solutions in banking or similar secured environments.GPU & Accelerated Compute: Deep, expert-level knowledge of NVIDIA GPU architectures, multi-node scaling (NVLink/NVSwitch), and high-performance networking (InfiniBand / RoCE v2).