Director of DevOps | 12+ Years | Ahmedabad Location
Vagaro · Ahmedabad, Gujarat, India
Apply & track with Apply EdgeAbout the Company :Vagaro is a global leader in cloud-based business management software, serving the salon, spa, beauty, fitness, and wellness industries. Our platform enables businesses to manage scheduling, payments, marketing, and operations from one centralized system.We are a rapidly growing, product-led company transforming how wellness businesses operate and connect with their customers.Why Join Us :5-day work week & flexible scheduleFun Fridays, Tech sessions & cultural eventsLearning resources & gaming zonesAnnual bonus & recognition programsLeave encashment & maternity benefits15 Paid Leaves + 11 HolidaysComprehensive family medical insurance Role: Director of Devops Location: Ahmedabad (Work from Office)Position Summary: We are looking for an experienced Director of DevOps, SRE & Infrastructure to lead the strategy, architecture, and operations of our Microsoft Azure Cloud Platform and engineering infrastructure. This role will own DevOps, SRE, Infrastructure, Platform Engineering, Observability, Security, Disaster Recovery, Scalability, and Cloud Cost Optimization. The ideal candidate combines strong technical expertise with proven leadership experience in designing and operating highly available, secure, scalable, and mission-critical SaaS platforms on Microsoft Azure. Strong hands-on and leadership experience with Microsoft Azure Cloud Platform is mandatory. Key Responsibilities: DevOps & Platform Engineering: Define and execute the DevOps and Platform Engineering strategy. Establish standardized CI/CD, release, deployment, and environment management practices. Drive automation, self-service infrastructure, and developer productivity. Establish engineering standards for Infrastructure as Code and configuration management. Promote safe deployment practices including Blue/Green, Canary, and Rolling deployments. Site Reliability Engineering: Build and mature the SRE organization and culture. Define and manage SLIs, SLOs, SLAs, and Error Budgets. Improve system availability, reliability, performance, and scalability. Establish production readiness and service ownership standards. Reduce operational toil and recurring production issues.Infrastructure & Architecture: Own the architecture and operation of Azure infrastructure. Design highly available, scalable, fault-tolerant, and multi-region systems. Establish standards for compute, networking, storage, databases, caching, messaging, and traffic management. Drive infrastructure modernization and cloud migration initiatives. Observability & Operations: Establish enterprise observability across metrics, logs, traces, APM, and infrastructure monitoring. Define monitoring, alerting, and operational dashboards. Establish effective incident management, escalation, and post-incident review processes. Drive reduction of MTTD, MTTR, and recurring incidents. Security & Compliance: Partner with Security and Engineering to establish secure infrastructure. Implement standards for IAM, secrets, certificates, encryption, network security, vulnerability management, and supply-chain security. Support SOC 2, ISO 27001, PCI DSS, and other applicable compliance requirements. Disaster Recovery & Business Continuity: Own infrastructure-level DR and Business Continuity strategy. Define and maintain RTO/RPO objectives. Establish backup, recovery, regional failover, and disaster recovery standards. Conduct regular DR and failover testing. Performance, Scalability & Cost: Establish infrastructure capacity and performance planning. Identify and eliminate infrastructure bottlenecks. Ensure platforms can scale with business growth. Own Azure cost optimization and FinOps initiatives. Balance reliability, performance, scalability, and infrastructure cost. Leadership: Lead DevOps, SRE, Infrastructure, and Platform Engineering teams. Recruit, mentor, and develop high-performing technical teams. Establish engineering objectives, standards, and KPIs. Partner with CTO/CPTO, Engineering, Architecture, Security, Product, QA, and Finance. Manage strategic technology vendors and infrastructure budgets. Core Technical Skill Set: Azure — Required: Strong production experience with: Azure App Services Azure Kubernetes Service (AKS) Azure Functions Azure Virtual Machines Azure Storage Azure SQL Cosmos DB Azure Cache for Redis Azure Service Bus Azure Event Hubs Azure Front Door Azure Application Gateway Azure Load Balancer Azure CDN Azure DNS Azure API Management Azure Virtual Network Azure Monitor Application Insights Log Analytics Azure Key Vault Microsoft Entra ID Azure Firewall Azure WAF Azure Backup / Site Recovery DevOps & CI/CD: GitHub / GitHub Enterprise Azure DevOps GitHub Actions CI/CD architecture Release automation Deployment automation GitOps Feature flags Blue/Green deployment Canary deployment Rolling deployment Infrastructure as Code: Terraform Bicep / ARM Ansible Helm Kubernetes manifests Policy as Code Containers & Kubernetes: Docker Kubernetes AKS Helm Kubernetes networking Autoscaling Container security Cluster management Kubernetes observability SRE: SLI / SLO / SLA Error Budgets Incident Management Production Readiness Reliability Engineering Capacity Planning Toil Reduction Chaos Engineering Fault Tolerance Disaster Recovery Observability: Azure Monitor Application Insights Log Analytics Prometheus Grafana Elastic OpenTelemetry Distributed Tracing APM Centralized Logging Alerting Networking: TCP/IP HTTP/HTTPS DNS TLS/SSL CDN WAF DDoS Protection Load Balancing API Gateway Reverse Proxy VNet Firewall VPN Routing Private Networking Security: Microsoft Entra ID IAM / RBAC OAuth 2.0 OpenID Connect Secrets Management Certificate Management Encryption / KMS Zero Trust Vulnerability Management Container Security SAST / DAST Supply-Chain Security Data & Messaging: Working knowledge of: SQL Server PostgreSQL MongoDB Redis Elasticsearch Azure Service Bus Azure Event Hubs Kafka RabbitMQ Distributed systems Replication Partitioning Sharding Backup & Recovery Leadership & Management Skills: Strategic planning Technical vision Engineering leadership Team building Hiring and talent development Executive communication Incident leadership Architecture decision-making Risk management Vendor management Budget management Cross-functional collaboration Stakeholder management Performance management Change management Experience & Qualifications: Required: 12+ years of experience in DevOps, SRE, Infrastructure, Cloud, Platform Engineering, or related fields. 5+ years of engineering leadership/management experience. Proven experience managing DevOps/SRE/Infrastructure teams. Strong hands-on experience with Microsoft Azure. Strong experience with Kubernetes/AKS and modern CI/CD. Experience operating large-scale production SaaS systems. Experience designing highly available and distributed systems. Experience with production incident management and disaster recovery. Strong understanding of security, observability, scalability, and cloud cost management. Preferred: Experience operating 24×7 mission-critical SaaS platforms. Experience with multi-region / active-active architectures. Experience with SOC 2, ISO 27001, PCI DSS, or similar compliance environments. Experience building or transforming DevOps/SRE organizations. Experience managing significant infrastructure/cloud budgets. Key Success Metrics: Reliability: Availability, SLO AchievementIncidents: MTTD, MTTR, Incident FrequencyDeployment: Deployment Frequency, Lead TimeQuality: Change Failure RateScalability: Capacity & PerformanceRecovery: RTO / RPOAutomation: Reduction in Operational ToilSecurity: Critical Vulnerability RemediationCost: Azure Cost OptimizationDR: Successful DR/Failover TestsObservability: Monitoring & Alert CoverageDeveloper Experience: CI/CD Reliability & Deployment EfficiencyIdeal Candidate: The ideal candidate is not just a DevOps manager. They should be able to operate at three levels: Strategic: Define the long-term DevOps, SRE, infrastructure, reliability, and platform strategy. Architectural: Make critical decisions around Azure, Kubernetes, networking, distributed systems, observability, security, scalability, and DR. Operational: Lead critical production incidents, improve reliability, and build systems that prevent recurring failures. Lead. Architect. Transform. Your journey as Vagaro’s Director of DevOps starts here—apply now!