Senior Site Reliability Engineer (SRE) – AWS & EKS
ZettaMine Labs Pvt. Ltd. · Bengaluru, Karnataka, India
Apply & track with Apply EdgeHello All,Greetings from ZettaMine!!!Role: Senior Site Reliability Engineer (SRE) – AWS & EKSExperience: 4+ Years & 7+ Years
This role requires strong expertise in resolving complex production incidents, building scalable automation and self-healing solutions, improving observability, and mentoring engineers.Key ResponsibilitiesLead troubleshooting and resolution of complex, high-severity production incidents across AWS infrastructure and Amazon EKS clusters.Act as an Incident Commander for major incidents and drive issues through complete resolution.Troubleshoot complex EKS/Kubernetes issues, including cluster-level failures, networking issues, autoscaling problems, performance bottlenecks, and upgrade-related incidents.Design and develop Python-based automation frameworks, internal tools, and operational solutions to eliminate manual effort and improve platform reliability.Architect, implement, and maintain Terraform-based Infrastructure as Code (IaC) with reusable, secure, and scalable modules and patterns.Design and implement self-healing automation and automated runbooks to detect known failure patterns and trigger remediation automatically.Drive the observability strategy by developing Grafana dashboards, monitoring solutions, and scalable alerting frameworks.Implement and maintain multi-region monitoring to provide consistent visibility into system health, latency, availability, and failover readiness.Perform advanced troubleshooting and root-cause analysis (RCA) using SQL and AWS CloudWatch Logs Insights.Own end-to-end ITSM Incident Management and Problem Management, including RCA, post-incident reviews, corrective actions, and long-term remediation.Proactively identify potential failure points, reliability risks, and performance bottlenecks before they impact production.Reduce operational workload by automating repetitive and recurring manual activities.Provide senior-level escalation support and participate in on-call rotations.Mentor junior and mid-level SRE/DevOps engineers in troubleshooting, automation, incident handling, and reliability engineering best practices.Partner with engineering, product, and leadership teams to clearly communicate incident impact, technical risks, root causes, and remediation plans.Mandatory CertificationsCandidates must have both of the following:AWS Certification, such as:AWS Certified Solutions Architect – ProfessionalAWS Certified DevOps Engineer – ProfessionalOr equivalent advanced AWS certificationKubernetes CertificationCertified Kubernetes Administrator (CKA)Or equivalent EKS/Kubernetes certificationInterested candidates share your updated CV to praneeth.n@zettamine.com or WhatsApp to 8977318915.Thanks& Regards,Praneeth