أبلاي إيدج ابدأ البحث عن عمل

Director of DevOps | 12+ Years | Ahmedabad Location

Vagaro · Ahmedabad, Gujarat, India

قدّم وتابع مع أبلاي إيدج
About the Company :Vagaro is a global leader in cloud-based business management software, serving the salon, spa, beauty, fitness, and wellness industries. Our platform enables businesses to manage scheduling, payments, marketing, and operations from one centralized system.We are a rapidly growing, product-led company transforming how wellness businesses operate and connect with their customers.Why Join Us :5-day work week & flexible scheduleFun Fridays, Tech sessions & cultural eventsLearning resources & gaming zonesAnnual bonus & recognition programsLeave encashment & maternity benefits15 Paid Leaves + 11 HolidaysComprehensive family medical insurance Role: Director of Devops Location: Ahmedabad (Work from Office)Position Summary: We are looking for an experienced Director of DevOps, SRE & Infrastructure to lead the strategy, architecture, and operations of our Microsoft Azure Cloud Platform and engineering infrastructure. This role will own DevOps, SRE, Infrastructure, Platform Engineering, Observability, Security, Disaster Recovery, Scalability, and Cloud Cost Optimization. The ideal candidate combines strong technical expertise with proven leadership experience in designing and operating highly available, secure, scalable, and mission-critical SaaS platforms on Microsoft Azure. Strong hands-on and leadership experience with Microsoft Azure Cloud Platform is mandatory.  Key Responsibilities: DevOps & Platform Engineering: Define and execute the DevOps and Platform Engineering strategy.  Establish standardized CI/CD, release, deployment, and environment management practices.  Drive automation, self-service infrastructure, and developer productivity.  Establish engineering standards for Infrastructure as Code and configuration management.  Promote safe deployment practices including Blue/Green, Canary, and Rolling deployments.  Site Reliability Engineering: Build and mature the SRE organization and culture.  Define and manage SLIs, SLOs, SLAs, and Error Budgets.  Improve system availability, reliability, performance, and scalability.  Establish production readiness and service ownership standards.  Reduce operational toil and recurring production issues.Infrastructure & Architecture: Own the architecture and operation of Azure infrastructure.  Design highly available, scalable, fault-tolerant, and multi-region systems.  Establish standards for compute, networking, storage, databases, caching, messaging, and traffic management.  Drive infrastructure modernization and cloud migration initiatives.  Observability & Operations: Establish enterprise observability across metrics, logs, traces, APM, and infrastructure monitoring.  Define monitoring, alerting, and operational dashboards.  Establish effective incident management, escalation, and post-incident review processes.  Drive reduction of MTTD, MTTR, and recurring incidents.  Security & Compliance: Partner with Security and Engineering to establish secure infrastructure.  Implement standards for IAM, secrets, certificates, encryption, network security, vulnerability management, and supply-chain security.  Support SOC 2, ISO 27001, PCI DSS, and other applicable compliance requirements.  Disaster Recovery & Business Continuity: Own infrastructure-level DR and Business Continuity strategy.  Define and maintain RTO/RPO objectives.  Establish backup, recovery, regional failover, and disaster recovery standards.  Conduct regular DR and failover testing.  Performance, Scalability & Cost: Establish infrastructure capacity and performance planning.  Identify and eliminate infrastructure bottlenecks.  Ensure platforms can scale with business growth.  Own Azure cost optimization and FinOps initiatives.  Balance reliability, performance, scalability, and infrastructure cost.  Leadership: Lead DevOps, SRE, Infrastructure, and Platform Engineering teams.  Recruit, mentor, and develop high-performing technical teams.  Establish engineering objectives, standards, and KPIs.  Partner with CTO/CPTO, Engineering, Architecture, Security, Product, QA, and Finance.  Manage strategic technology vendors and infrastructure budgets.  Core Technical Skill Set: Azure — Required: Strong production experience with: Azure App Services  Azure Kubernetes Service (AKS)  Azure Functions  Azure Virtual Machines  Azure Storage  Azure SQL  Cosmos DB  Azure Cache for Redis  Azure Service Bus  Azure Event Hubs  Azure Front Door  Azure Application Gateway  Azure Load Balancer  Azure CDN  Azure DNS  Azure API Management  Azure Virtual Network  Azure Monitor  Application Insights  Log Analytics  Azure Key Vault  Microsoft Entra ID  Azure Firewall  Azure WAF  Azure Backup / Site Recovery   DevOps & CI/CD: GitHub / GitHub Enterprise  Azure DevOps  GitHub Actions  CI/CD architecture  Release automation  Deployment automation  GitOps  Feature flags  Blue/Green deployment  Canary deployment  Rolling deployment  Infrastructure as Code: Terraform  Bicep / ARM  Ansible  Helm  Kubernetes manifests  Policy as Code Containers & Kubernetes: Docker  Kubernetes  AKS  Helm  Kubernetes networking  Autoscaling  Container security  Cluster management  Kubernetes observability  SRE: SLI / SLO / SLA  Error Budgets  Incident Management  Production Readiness  Reliability Engineering  Capacity Planning  Toil Reduction  Chaos Engineering  Fault Tolerance  Disaster Recovery  Observability: Azure Monitor  Application Insights  Log Analytics  Prometheus  Grafana  Elastic  OpenTelemetry  Distributed Tracing  APM  Centralized Logging  Alerting  Networking: TCP/IP  HTTP/HTTPS  DNS  TLS/SSL  CDN  WAF  DDoS Protection  Load Balancing  API Gateway  Reverse Proxy  VNet  Firewall  VPN  Routing  Private Networking  Security: Microsoft Entra ID  IAM / RBAC  OAuth 2.0  OpenID Connect  Secrets Management  Certificate Management  Encryption / KMS  Zero Trust  Vulnerability Management  Container Security  SAST / DAST  Supply-Chain Security  Data & Messaging: Working knowledge of: SQL Server  PostgreSQL  MongoDB  Redis  Elasticsearch  Azure Service Bus  Azure Event Hubs  Kafka  RabbitMQ  Distributed systems  Replication  Partitioning  Sharding  Backup & Recovery   Leadership & Management Skills: Strategic planning  Technical vision  Engineering leadership  Team building  Hiring and talent development  Executive communication  Incident leadership  Architecture decision-making  Risk management  Vendor management  Budget management  Cross-functional collaboration  Stakeholder management  Performance management  Change management  Experience & Qualifications: Required: 12+ years of experience in DevOps, SRE, Infrastructure, Cloud, Platform Engineering, or related fields.  5+ years of engineering leadership/management experience.  Proven experience managing DevOps/SRE/Infrastructure teams.  Strong hands-on experience with Microsoft Azure.  Strong experience with Kubernetes/AKS and modern CI/CD.  Experience operating large-scale production SaaS systems.  Experience designing highly available and distributed systems.  Experience with production incident management and disaster recovery.  Strong understanding of security, observability, scalability, and cloud cost management.  Preferred: Experience operating 24×7 mission-critical SaaS platforms.  Experience with multi-region / active-active architectures.  Experience with SOC 2, ISO 27001, PCI DSS, or similar compliance environments.  Experience building or transforming DevOps/SRE organizations.  Experience managing significant infrastructure/cloud budgets.   Key Success Metrics: Reliability: Availability, SLO AchievementIncidents: MTTD, MTTR, Incident FrequencyDeployment: Deployment Frequency, Lead TimeQuality: Change Failure RateScalability: Capacity & PerformanceRecovery: RTO / RPOAutomation: Reduction in Operational ToilSecurity: Critical Vulnerability RemediationCost: Azure Cost OptimizationDR: Successful DR/Failover TestsObservability: Monitoring & Alert CoverageDeveloper Experience: CI/CD Reliability & Deployment EfficiencyIdeal Candidate: The ideal candidate is not just a DevOps manager. They should be able to operate at three levels: Strategic: Define the long-term DevOps, SRE, infrastructure, reliability, and platform strategy. Architectural: Make critical decisions around Azure, Kubernetes, networking, distributed systems, observability, security, scalability, and DR. Operational: Lead critical production incidents, improve reliability, and build systems that prevent recurring failures.  Lead. Architect. Transform. Your journey as Vagaro’s Director of DevOps starts here—apply now!