Director – Site Reliability Engineering (SRE)
Umanist NA · Navi Mumbai, Maharashtra, India
Apply & track with Apply EdgeEstablish and drive reliability engineering strategy across the organization. Define and manage SLIs, SLOs, SLAs, error budgets, and reliability KPIs. Drive observability, monitoring, alerting, and proactive performance management. Establish and improve incident response, escalation, troubleshooting, and post-incident review processes. Partner with Software Engineering, Product, Cloud/Platform, Security, and Infrastructure teams. Apply software engineering principles to automate and solve reliability and operational challenges. Improve system availability, scalability, resilience, performance, and operational efficiency. Drive CI/CD improvements and deployment reliability. Architect and support applications and infrastructure running on high-growth cloud platforms. Lead reliability programs across multiple B2B SaaS products. Establish engineering metrics and use data/KPIs to identify and resolve operational gaps. Drive automation and reduction of repetitive operational work. Mentor engineering leaders and senior SRE engineers. Communicate reliability strategy, risks, metrics, and business impact to senior leadership. Influence engineering teams and cross-functional stakeholders to adopt reliability best practices. Bachelor's degree in Computer Science, Information Systems, Engineering, or a related field.18+ years of experience in Software Engineering, SRE, Site Reliability, Platform Engineering, or related reliability/engineering roles. 5+ years of leadership experience at Director level or equivalent. Strong experience in SaaS / B2B SaaS / cloud product companies. Proven ability to apply software engineering principles and practices to solve reliability and operational challenges. Strong expertise in SLI/SLO, monitoring, observability, and reliability engineering. Strong experience with CI/CD and modern software delivery practices. Strong experience with incident response, problem management, RCA, and production operations. Strong AWS expertise. Experience with container orchestration, such as Kubernetes. Experience leading reliability programs across multiple SaaS products. Experience architecting applications or infrastructure for high-growth cloud platforms. Experience with large-scale distributed systems in B2B SaaS environments. Strong leadership, communication, stakeholder management, and influencing skills. Demonstrated experience driving operational excellence through metrics, KPIs, SLOs, and reliability objectives. Preferred SkillsKubernetes / containerized environments AWS cloud architecture Infrastructure and application observability Distributed systems Microservices Infrastructure automation Infrastructure-as-Code Terraform Prometheus / Grafana Datadog / New Relic / Splunk or similar observability platforms CI/CD platforms Disaster recovery and business continuity Capacity and performance engineering Chaos/resilience engineering Skills: kubernetes,large-scale distributed systems,metrics,container orchestration,production operations,kpis,cloud,multiple saas products,problem management,sli/slo,leadership,ci/cd,b2b saas environments,observability,slos,aws,saas,b2b,monitoring,reliability engineering