Apply Edge Start your job search

Senior Site Reliability Engineer (SRE) – AWS & EKS

ZettaMine Labs Pvt. Ltd. · Bengaluru, Karnataka, India

Apply & track with Apply Edge

Hello All,Greetings from ZettaMine!!!Role: Senior Site Reliability Engineer (SRE) – AWS & EKSExperience: 4+ Years & 7+ Years

Location: Bangalore, Pune, ChennaiNotice Period: Immediate JoinersShift Timing: 2:00 PM – 10:00 PMRole: Senior Site Reliability EngineerPrimary Skills: AWS, Amazon EKS, Kubernetes, Python, TerraformMandatory Certifications: AWS Certification + CKA/Kubernetes CertificationHiring for one of the Big 4's ClientsWe are looking for Immediate to 15 Days Joiners onlyJob SummaryWe are looking for an experienced Senior Site Reliability Engineer (SRE) with deep hands-on expertise in AWS cloud infrastructure, Amazon EKS/Kubernetes, Python automation, Terraform, observability, and production incident management.The ideal candidate will take ownership of reliability, availability, performance, and operational excellence across cloud infrastructure and Kubernetes platforms.

This role requires strong expertise in resolving complex production incidents, building scalable automation and self-healing solutions, improving observability, and mentoring engineers.Key ResponsibilitiesLead troubleshooting and resolution of complex, high-severity production incidents across AWS infrastructure and Amazon EKS clusters.Act as an Incident Commander for major incidents and drive issues through complete resolution.Troubleshoot complex EKS/Kubernetes issues, including cluster-level failures, networking issues, autoscaling problems, performance bottlenecks, and upgrade-related incidents.Design and develop Python-based automation frameworks, internal tools, and operational solutions to eliminate manual effort and improve platform reliability.Architect, implement, and maintain Terraform-based Infrastructure as Code (IaC) with reusable, secure, and scalable modules and patterns.Design and implement self-healing automation and automated runbooks to detect known failure patterns and trigger remediation automatically.Drive the observability strategy by developing Grafana dashboards, monitoring solutions, and scalable alerting frameworks.Implement and maintain multi-region monitoring to provide consistent visibility into system health, latency, availability, and failover readiness.Perform advanced troubleshooting and root-cause analysis (RCA) using SQL and AWS CloudWatch Logs Insights.Own end-to-end ITSM Incident Management and Problem Management, including RCA, post-incident reviews, corrective actions, and long-term remediation.Proactively identify potential failure points, reliability risks, and performance bottlenecks before they impact production.Reduce operational workload by automating repetitive and recurring manual activities.Provide senior-level escalation support and participate in on-call rotations.Mentor junior and mid-level SRE/DevOps engineers in troubleshooting, automation, incident handling, and reliability engineering best practices.Partner with engineering, product, and leadership teams to clearly communicate incident impact, technical risks, root causes, and remediation plans.Mandatory CertificationsCandidates must have both of the following:AWS Certification, such as:AWS Certified Solutions Architect – ProfessionalAWS Certified DevOps Engineer – ProfessionalOr equivalent advanced AWS certificationKubernetes CertificationCertified Kubernetes Administrator (CKA)Or equivalent EKS/Kubernetes certificationInterested candidates share your updated CV to praneeth.n@zettamine.com or WhatsApp to 8977318915.Thanks& Regards,Praneeth