Apply Edge Start your job search

Staff TechOps & Support Engineer

Sumerge · Cairo, Cairo, Egypt

Apply & track with Apply Edge

Our Staff TechOps & Support Engineer Ensures the stability, performance, and availability of production systems by managing infrastructure, monitoring environments, responding to incidents, automating operational tasks, supporting deployments, and maintaining security, backups, and documentation ResponsibilitiesEnsure the availability, performance, and resilience of production environments by proactively monitoring systems and resolving complex operational issuesAdminister and optimize CI/CD pipelines, containerized platforms, IBM Cloud Pak Stacks, Confluent, Elasticsearch, and supporting infrastructure to improve operational efficiencyLead incident response, perform root cause analysis, and implement preventive actions to reduce recurring issues and improve service reliabilityDevelop and enhance monitoring, logging, alerting, automation, and operational processes to improve system observability and reduce manual effortProvide technical leadership and mentorship to junior engineers while supporting deployments, release management, and production readiness activitiesMaintain system security, backup, disaster recovery, and compliance standards while ensuring accurate operational documentation and runbooksCollaborate with Development, Platform Engineering, QUALITY, Network, and Security teams to resolve complex technical challenges and continuously improve production operationsRequirementsBachelor's degree or Diploma in Computer Science, Engineering, or a related field4-6 years of experience in Technical Operations, Production Support, Site Reliability Engineering (SRE), DevOps, or System AdministrationStrong experience with Linux/Windows administration, Docker, Kubernetes, OpenShift, CI/CD pipelines, automation tools, scripting (Bash, PowerShell, Python), and configuration managementHands-on experience with IBM Cloud Pak Stacks (CP4BA, CP4I, CP4D), Confluent, Elasticsearch, Databases (SQL, DB2, MongoDB, etc.), JVM troubleshooting, monitoring, and observability platformsStrong understanding of networking, security best practices, IAM, production support, backup, disaster recovery, scalability, performance tuning, and service reliability principlesProven experience in incident management, troubleshooting, root cause analysis, and implementing operational improvements within enterprise production environmentsExcellent technical leadership, analytical, communication, mentoring, and collaboration skills with the ability to manage multiple priorities effectively