أبلاي إيدج ابدأ البحث عن عمل

Site Reliability Engineer

Throne Solutions · Riyadh, Riyadh, Saudi Arabia

قدّم وتابع مع أبلاي إيدج
Job Title: Site Reliability Engineer (SRE) – Freelance ContractCompany: Throne SolutionsLocation: Riyadh, Saudi ArabiaEmployment Type: Freelance ContractExperience Required: 5–8 YearsAbout Throne SolutionsThrone Solutions is seeking an experienced and motivated Site Reliability Engineer (SRE) to join our team on a freelance contract basis in Riyadh, Saudi Arabia. The successful candidate will be responsible for ensuring the availability, scalability, performance, and reliability of enterprise production environments through automation, cloud-native technologies, proactive monitoring, and operational excellence. This role requires strong expertise in AWS cloud infrastructure, Kubernetes, CI/CD, and incident management within enterprise production environments.Role SummaryAs a Site Reliability Engineer, you will bridge software engineering and IT operations by designing resilient cloud infrastructure, automating operational processes, improving system reliability, and minimizing downtime. You will collaborate closely with development, infrastructure, networking, and security teams to maintain highly available production systems while driving continuous improvement through automation, observability, and operational best practices.Key ResponsibilitiesDesign, build, and maintain highly available, scalable, and secure AWS cloud infrastructure.Provision and manage cloud resources using Infrastructure as Code (IaC) tools such as Terraform and AWS CloudFormation.Deploy, administer, and optimize Kubernetes clusters and Docker-based containerized applications.Develop automation scripts using Python, Bash, or Go to eliminate manual operational tasks and improve operational efficiency.Design, implement, and maintain CI/CD pipelines using Jenkins, GitLab CI/CD, or similar DevOps platforms.Implement and manage monitoring, logging, and observability solutions using Prometheus, Grafana, Splunk, Datadog, AWS CloudWatch, or equivalent platforms.Monitor application health, infrastructure performance, and service availability to proactively identify and resolve issues.Lead production incident response, perform Root Cause Analysis (RCA), and implement corrective and preventive actions.Define and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Error Budgets to improve platform reliability.Participate in a 24×7 on-call rotation and provide production support for business-critical applications and infrastructure.Improve system performance, scalability, reliability, and Mean Time to Recovery (MTTR) through automation and continuous optimization.Develop and maintain operational documentation, runbooks, disaster recovery procedures, and standard operating procedures.Collaborate with Development, DevOps, Infrastructure, Security, and Network teams to support deployments and operational readiness.Implement security best practices across cloud infrastructure, Kubernetes clusters, and CI/CD pipelines.Ensure compliance with ITIL Incident, Problem, Change, and Release Management processes.Support enterprise production environments, including Cisco-based infrastructure where applicable.Required QualificationsBachelor's degree in Computer Science, Information Technology, Software Engineering, Computer Engineering, or a related discipline.5–8 years of professional experience in Site Reliability Engineering (SRE), DevOps, Cloud Engineering, or Production Support.Proven experience supporting mission-critical enterprise production environments.Must be authorized to work in Saudi Arabia.Mandatory Technical SkillsCloud PlatformsAmazon Web Services (AWS)EC2VPCIAMRDSS3ELBAuto ScalingRoute 53AWS CloudWatchInfrastructure as Code (IaC)TerraformAWS CloudFormationContainerization & OrchestrationKubernetesDockerOperating SystemsLinux Administration (Red Hat, CentOS, Ubuntu)Programming & AutomationPythonBashGo (Preferred)CI/CD & DevOpsJenkinsGitLab CI/CDGitGitHubMonitoring & ObservabilityPrometheusGrafanaSplunkDatadogAWS CloudWatchIncident & ITSM ToolsServiceNowJiraITIL-based Service ManagementNetworking FundamentalsTCP/IPDNSHTTP/HTTPSLoad BalancingVPNFirewallsBasic Cisco NetworkingPreferred SkillsExperience working in Cisco enterprise environments.Knowledge of cloud security tools and security best practices.Experience with Infrastructure Monitoring and Application Performance Monitoring (APM).Familiarity with container security, Kubernetes security, and DevSecOps practices.Experience with Helm, ArgoCD, or GitOps methodologies.Strong understanding of microservices architecture and distributed systems.Experience with multi-cloud or hybrid cloud environments.Preferred CertificationsAWS Certified Solutions Architect – Associate or ProfessionalAWS Certified DevOps Engineer – ProfessionalCertified Kubernetes Administrator (CKA)Certified Kubernetes Application Developer (CKAD)HashiCorp Terraform AssociateRed Hat Certified System Administrator (RHCSA)ITIL Foundation CertificationKey Performance OutcomesMaintain high availability and reliability of enterprise production systems.Improve overall service uptime and platform resilience.Reduce Mean Time to Detect (MTTD) and Mean Time to Recovery (MTTR).Increase operational efficiency through automation and infrastructure optimization.Enhance monitoring, observability, and incident response capabilities.Deliver secure, scalable, resilient, and cost-effective cloud infrastructure.Ensure compliance with SLAs, SLOs, and operational best practices.Required CompetenciesStrong analytical and problem-solving skills.Excellent troubleshooting abilities in complex production environments.Strong communication and stakeholder management skills.Ability to perform effectively under pressure during critical production incidents.Automation-first mindset with a focus on continuous improvement.Excellent documentation and technical communication skills.Strong collaboration across development, infrastructure, networking, and security teams.Self-motivated with a passion for learning and adopting emerging cloud technologies.Mandatory RequirementsBachelor's Degree in Computer Science, Information Technology, Software Engineering, Computer Engineering, or a related field.Authorized to work in Saudi Arabia.Availability to work onsite in Riyadh.Minimum 5 years of hands-on experience in Site Reliability Engineering (SRE), DevOps, Cloud Engineering, or Production Support.Experience supporting AWS cloud environments, Kubernetes, and enterprise production systems.Willingness to participate in on-call support and shift rotations as required.