Apply Edge Start your job search

Senior Manager - Site Reliability Engineer|NR-2026-0246

Media.net · Bengaluru, Karnataka, India

Apply & track with Apply Edge
Senior Manager - Site Reliability EngineerAbout the RoleWe are looking for a Senior Manager - Site Reliability Engineering to lead a team of 8+ Site Reliability Engineers within our core platform engineering organization. The team supports a high-throughput, low-latency distributed platform that processes more than 500 billion requests daily, with stringent performance requirements where every millisecond matters.As a leader, you will own the reliability charter for this platform: setting technical direction, growing engineers, partnering with product and development leadership, and ensuring the infrastructure is resilient, scalable, and always available.The team is also evolving rapidly toward agentic AI. You should be well-versed in the modern AI ecosystem and capable of driving the adoption of AI-powered tools, workflows, and automation frameworks that improve engineering productivity, operational excellence, and platform reliability.What You'll Do People Leadership & Team Management • Lead, mentor, and grow a team of 8-10 SREs across varying experience levels, fostering a culture of ownership, learning, and psychological safety.• Drive AI adoption, AIOps, intelligent automation, and agentic AI initiatives. • Conduct regular 1:1s, performance reviews, and career development conversations; define clear growth paths for individual contributors.• Drive hiring: partner with recruiting to define role requirements, lead interviews, and build a diverse, highperforming SRE team.• Manage team workload, priorities, and on-call rotations to prevent burnout while maintaining operational excellence.• Champion SRE best practices, including error budgets, SLOs/SLIs/SLAs, and blameless post-mortems, and instill a data-driven reliability cultureInfrastructure Strategy & Ownership • Own the end-to-end reliability roadmap for a high-scale, real-time platform, including load balancers, data stores, CI/CD pipelines, observability stacks, and networking layers.• Define and drive multi-quarter infrastructure programs focused on resilience, scalability, performance, and cost efficiency to support more than 500 billion requests daily.• Develop and enforce policies and procedures that improve platform stability; represent the team in architectural reviews and cross-functional planning.• Establish infrastructure-as-code standards using tools such as Terraform, Puppet, or Ansible and ensure consistent adoption across the team.• Bring deep, hands-on technical expertise in Kubernetes, including cluster design, multi-tenancy, networking, autoscaling, upgrades, and operational best practices at scale.• Drive architectural ownership across the SRE landscape, making high-judgment decisions on platform design, service topology, and infrastructure trade-offs.• Define and evolve the cloud and container strategy, including Kubernetes, service mesh, CI/CD, and observability, to support rapid product growth and engineering velocity.• Lead capacity planning, performance engineering, and resilience initiatives, including disaster recovery, chaos engineering, and multi-region availability.Cross-Functional Collaboration • Partner closely with software engineering leads to set quality and performance benchmarks; ensure production-readiness criteria are met before every deployment.• Participate in system design reviews, providing leadership-level input on infrastructure reliability, security, and operability.• Serve as the primary point of escalation for high-severity incidents; coordinate cross-team response and ensure transparent communication to stakeholders.• Collaborate with product management to balance feature velocity against reliability investments, using error budgets as a shared language.Tooling, Automation & Performance • Drive the vision for internal SRE tooling, including monitoring, alerting, deployment automation, failure detection, and self-healing systems.• Champion full-stack performance optimization, from request handling and network paths to database query tuning, with a focus on sub-millisecond latency targets.• Ensure robust observability across all services using Prometheus, Grafana, or the ELK stack; establish dashboards and alerts that provide real-time operational insight.• Oversee CI/CD pipeline health using tools such as Jenkins and Argo CD, and continuously improve deployment speed, safety, and rollback capabilities.Who Should Apply Required Qualifications• B.Tech, M.Tech, or equivalent degree in Computer Science, Information Technology, or a related field.• 13+ years of overall experience in SRE, infrastructure engineering, or DevOps, including at least 2 years in a people management role leading teams of 8+ engineers.• Proven track record of building and scaling SRE teams supporting high-volume, latency-sensitive distributed systems.• Strong technical foundation in networking concepts, including TCP/IP, routing, and SDN, as well as modern software architectures.• Hands-on proficiency in Python, Go, or Ruby, with a strong automation mindset. • Deep experience with Kubernetes and container orchestration; familiarity with GCP or AWS, in addition to operating large-scale co-location data centers.• Strong understanding of agentic AI, AIOps, MCP, A2A, and AI automation frameworks.• Demonstrated ability to define and operate against SLOs, error budgets, and reliability metrics at scale.• Ability to independently own complex problem statements, set team priorities, and drive solutions end to end.Preferred Skills & Tool Expertise • Infrastructure as Code: Terraform, Puppet, or Ansible.• Monitoring & Logging: Prometheus, Grafana, and the ELK stack. • CI/CD Pipelines: Jenkins and Argo CD.• Databases: MySQL, Redis, Aerospike, HBase, or similar technologies.• Web Proxies & Service Networking: Envoy and Nginx.• Version Control: Git-based workflows at team scale.• Experience operating high-throughput, low-latency platforms or large-scale real-time distributed systems is a strong plus.• Familiarity with incident management frameworks, PagerDuty ecosystems, and a blameless post-mortem culture.