Sr DevOps Engineer
Boubyan Digital Factory · New Cairo, Cairo, Egypt
Apply & track with Apply EdgeCore Technology StackAWS (5+ years) •EKS • ECS • Terraform• Istio •Datadog • GitLab CI/CD (Advanced) •ArgoCD • Serverless Framework• Node.js •Python • Temporal• Docker •CloudFormationRole OverviewAs a Senior DevOps Engineer, you will take a hands-on approach to the delivery and automation of our cloud infrastructure, supporting continuous integration and deployment. You will work closely with the DevOps Technical Lead and collaborate with software engineering teams to build, secure, and scale our platform.This is a financial services environment where security, compliance, and reliability are not optional. You will be expected to build and maintain infrastructure that meets SOC 2 Type II, PCI DSS, and GDPR as a basic requirement while keeping the developer experience fast and friction-free – partnering with engineering teams.ResponsibilitiesInfrastructure & PlatformDesign, build, and maintain cloud-native infrastructure on AWS that is secure, automated, scalable, and highly available.Own and evolve our Kubernetes platform (EKS) including cluster operations, upgrades, networking, and security hardening.Manage and operate ECS (Fargate and EC2 launch types) for containerised workloads, including task definitions, service scaling, and deployment strategies.Design and maintain advanced AWS networking: Transit Gateways for multi-VPC/multi-account connectivity, Site-to-Site VPN, AWS Client VPN, VPC peering, Direct Connect, Route 53 DNS resolution, and network segmentation across environments.Manage and advance our Istio service mesh for traffic management, mTLS enforcement, canary deployments, and observability across microservices.Provision, configure, and maintain all AWS infrastructure as code through Terraform, enforcing best practices around modules, state management, and drift detection.Maintain and extend CloudFormation stacks where applicable, including nested stack architectures and cross-stack dependencies.Implement and manage GitOps workflows using ArgoCD for declarative, auditable Kubernetes deployments.Build and operate serverless workloads using the Serverless Framework for event-driven and API-backed services.Maintain and support the infrastructure around Temporal: cluster operations, namespace management, monitoring, scaling, and ensuring high availability for engineering teams building durable workflows.Observability & ReliabilityDeploy and configure Datadog across the full stack: infrastructure metrics, APM (distributed tracing), log management, synthetic monitoring, and SLOs.Design tagging strategies and dashboard standards that give engineering teams self-service visibility into their services.Build meaningful alerting with low noise: composite monitors, anomaly detection, and error budget burn-rate alerts.Perform infrastructure cost analysis and optimization on an ongoing basis, including reserved capacity planning and right-sizing.Security & ComplianceMaintain infrastructure that satisfies SOC 2 Type II, PCI DSS, and GDPR compliance requirements.Implement and enforce encryption at rest and in transit across all services using AWS KMS customer-managed keys.Ensure comprehensive audit logging: CloudTrail organization trails, application-level audit logs, EKS audit logs, and tamper-proof log retention.Enforce least-privilege IAM policies, SCPs across AWS Organizations, and secret management through AWS Secrets Manager or SSM Parameter Store.Maintain network segmentation and security controls for cardholder data environments in line with PCI DSS.Collaborate with the AppSec engineer on vulnerability management, secret scanning in CI, and compliance evidence collection for audits.CI/CD & Developer ExperienceOwn and architect advanced GitLab CI/CD pipelines across all services: multi-stage, multi-environment pipelines with dynamic child pipelines, DAG dependencies, and rules-based execution.Enforce Principle of Least Privilege (POLP) across CI/CD: scoped CI/CD variables, protected environments, least-privilege runner tokens, per-job service accounts, and minimal container image permissions.Design and maintain .gitlab-ci.yml standards across the organisation: reusable includes, extends, component templates, and shared CI libraries to eliminate duplication and enforce consistency.Build and maintain deployment tooling for Node.js and Python services across containerized and serverless targets.Ensure pipeline compliance: MR-based approvals, signed artifacts, deployment gates, approval rules, and full traceability from ticket to production.Manage GitLab Runner infrastructure: autoscaling runners on AWS (EC2/EKS), runner caching strategies, Docker-in-Docker vs Kaniko for container builds, and runner security hardening.Optimise pipeline performance: caching, artifacts, parallel jobs, resource groups, and pipeline analytics to keep feedback loops fast.Collaboration & On-CallWork collaboratively with software engineering teams to define infrastructure, deployment, and operational requirements.Contribute to the evaluation of new technologies and vendor products that drive additional value across the stack.Share knowledge and mentor junior team members through code reviews, documentation, and pairing sessions.Translate non-technical business requirements into technical infrastructure requirements.Participate in an on-call rotation: respond to production incidents, perform root cause analysis, and drive post-incident improvements to prevent recurrence.Required Skills & ExperienceMust Have (in priority order)Deep AWS expertise across core services: VPC, IAM, EKS, ECS (Fargate and EC2), Lambda, S3, RDS, KMS, CloudTrail, Route 53, API Gateway, and Organizations. (5+ years)Advanced AWS networking: Transit Gateway architectures, Site-to-Site VPN, AWS Client VPN, VPC design and subnetting, DNS (Route 53 private/public hosted zones, Resolver endpoints), VPC peering, Direct Connect, PrivateLink, and load balancing (ALB/NLB).Production Kubernetes experience (EKS strongly preferred) including cluster operations, RBAC, networking, HPA, and troubleshooting. (3+ years)Hands-on Istio service mesh experience: traffic management, mTLS, VirtualService/DestinationRule configuration, sidecar tuning, and debugging with istioctl.(3+ years)Advanced Terraform experience including modules, workspaces, state management, and CI-driven plan/apply workflows. (3+ years)Advanced GitLab CI/CD expertise: excellent understanding of CI syntax (rules, needs, DAG, dynamic child pipelines, includes/extends, components), pipeline security (POLP, protected variables/environments, runner isolation), MR approval workflows, and runner infrastructure management. (4–5 years)Serverless Framework for deploying and managing Lambda-based workloads.Datadog implementation and operations: Agent deployment in Kubernetes, APM instrumentation, log pipelines, custom metrics, dashboards, monitors, and SLOs.Security and compliance mindset: practical experience with SOC 2, PCI DSS, or similar regulatory frameworks in a fintech or financial services environment.Node.js hands-on experience: comfortable reading, debugging, and writing Node.js code. Understanding of the runtime, LTS lifecycle, and packaging.Python scripting for automation: boto3, CLI tools, CI scripts, data processing. Comfortable with concurrency patterns and packaging.ArgoCD for GitOps-based Kubernetes deployments.Docker and container best practices: image optimization, multi-stage builds, vulnerability scanning.Required CertificationsAWS Certified DevOps Engineer — ProfessionalAWS Certified Solutions Architect — AssociateHashiCorp Certified: Terraform Associate or EngineerNice to HaveTemporal platform operations: cluster deployment, namespace configuration, visibility store tuning, and worker infrastructure.CloudFormation including nested stacks, macros, and cross-stack references.Certified Kubernetes Security Specialist (CKS).Azure Active Directory and Microsoft 365 integration.Experience with JIRA and Confluence for project tracking and documentation.Familiarity with SIEM toolsets and cloud log management platforms beyond Datadog.Experience running SOC 2 Type II audit cycles end-to-end, including evidence collection from the DevOps side.Contributions to open-source infrastructure tooling.What We ValueOwnership: you see a problem, you fix it. You don't wait for a ticket.Pragmatism over perfection: ship secure, maintainable solutions, iterate later.Clear communication: you can explain a VPC design to a product manager and a compliance requirement to an engineer.Curiosity: the appetite to learn and evaluate a wide variety of open-source technologies and tools.Reliability: comfortable with on-call responsibilities and committed to keeping production healthy.