Systems Engineer - Infrastructure & Disaster Recovery (National Talent)
Dubai Department of Economy and Tourism · Dubai, United Arab Emirates
قدّم وتابع مع أبلاي إيدجJOB OBJECTIVE:The Systems Engineer, Infrastructure & Disaster Recovery is responsible for maintaining, optimising and recovering the organisation’s server, virtualisation, container, backup and selected cloud infrastructure.The Systems Engineer ensures that infrastructure services remain available, secure, supportable and recoverable through effective platform administration, monitoring, automation, backup management and disaster-recovery practices.The Systems Engineer implements approved infrastructure standards and changes, resolves complex technical incidents and works with Network, Database, Applications, DevOps and Cybersecurity teams to maintain reliable infrastructure services.Core Functional Responsibilities & Subject Matter ExpertiseInfrastructure Platform AdministrationAdminister and maintain supported Windows Server and enterprise Linux environments.Provision, configure, patch, upgrade and decommission physical and virtual servers in accordance with approved standards.Administer assigned infrastructure services, including Active Directory, Group Policy, DNS, DHCP and file services.Maintain server configuration, security-hardening and supported-version standards.Monitor infrastructure availability, performance, capacity and resource utilisation.Investigate and resolve complex operating-system and infrastructure incidents.Support server migration, consolidation, lifecycle renewal and technology refresh activities.Virtualisation and Container InfrastructureAdminister and maintain the organisation’s approved virtualisation platforms and clusters.Provision, resize, migrate and decommission virtual machines.Maintain high availability, resource balancing and live-migration capabilities.Monitor host, cluster and virtual-machine capacity, performance and availability.Administer approved Kubernetes infrastructure, including cluster and node lifecycle, access controls, certificates and platform health.Maintain Kubernetes high availability, infrastructure-level backup and cluster-restoration capabilities.Troubleshoot control-plane, node, scheduling, capacity and cluster-level incidents.Coordinate application deployments and workload-specific requirements with DevOps and Application teams.Backup, Restoration and Data ProtectionAdminister the organisation’s enterprise backup and recovery platform.Maintain backup protection for servers, virtual machines, file systems, infrastructure services and assigned platform components.Maintain backup schedules, retention requirements, storage policies and recovery copies.Monitor backup, replication and recovery activities and resolve failed or incomplete jobs.Conduct scheduled restoration tests to confirm the integrity and recoverability of protected infrastructure.Maintain secure, immutable or isolated recovery copies in accordance with approved requirements.Identify backup-coverage, capacity or recoverability gaps and coordinate corrective actions.Maintain current backup and restoration procedures and reports.Infrastructure Resilience and Disaster RecoveryMaintain systems-domain disaster-recovery procedures, recovery sequences and technical runbooks.Ensure infrastructure recovery arrangements support approved recovery priorities, Recovery Time Objectives and Recovery Point Objectives.Identify infrastructure dependencies involving identity, networking, storage, databases, security services and applications.Maintain recovery procedures for servers, virtual machines, infrastructure services, Kubernetes platforms and selected cloud workloads.Plan and execute scheduled restoration, failover, failback and disaster-recovery exercises.Validate infrastructure availability, connectivity and platform health following recovery.Record recovery results, achieved recovery times, failed activities and technical gaps.Recommend resilience and recovery improvements and track approved actions through to closure.Maintain disaster-recovery environments in an operational, patched and recoverable state.Ensure relevant production changes are reflected in recovery configurations, capacity requirements and runbooks.Cloud Infrastructure and AutomationAdminister selected cloud infrastructure services within the approved operational scope.Maintain approved cloud virtual machines, storage and infrastructure-recovery services.Conduct approved cloud recovery tests, failovers, reprotection and failback activities.Monitor cloud infrastructure availability, performance and resource utilisation.Automate routine infrastructure provisioning, configuration, patching, monitoring and recovery activities.Develop and maintain controlled automation using approved scripting and infrastructure-as-code tools.Maintain scripts, templates and configuration files in approved version-control repositories.Apply testing, peer review, security validation and change control to automated infrastructure changes.Identify configuration drift and repetitive administrative activities requiring automation.The role does not own enterprise cloud architecture, application deployment pipelines or application-release automation.