Director of Infrastructure & Reliability
PracticeSuite, Inc. · Tampa, FL
Apply & track with Apply EdgeLocation: Tampa, FLReports to: VP Engineering Department: Engineering & System OperationsAbout the RoleOpportunity to lead Infrastructure and site reliability engineering with authority around DevOps as well. You'll take a platform that's been reliably running critical healthcare operations and evolve it into a modern, resilient, forward-looking infrastructure organization with the autonomy to set the technical standards and the roadmap from day one.This is a founding leadership role: you'll make your first hire, help that person build their own team, and establish the operating model that the entire infrastructure function will run on for years to come. If you've always wanted to build an SRE function the right way with executive backing, room to define what "good" looks like, and a real chance to leave your mark on how a company operates this is that opportunity.ResponsibilitiesStand up SRE as a function: incident command, US-hours on-call coverage, post-incident reviews, and SLOsSet infrastructure direction: across current servers, and new solutions including vendor recommendationsDrive Cloud Infrastructure: Design and implement AWS accounts, VPC, IAM, and data-store shape. Own cost, scaling, and resilience trade-offs.Own CI/CD and operational foundations: environment and domain management, observability, secrets management, access control, and backup/restore evidenceDefine the operating split between Server Admin (running the current estate) and SRE (reliability, toil reduction, and the new platform)Requirements3+ years leading platform and/or infrastructure teamsBuilt or led an SRE, platform, or infrastructure team at Director or Manager level, and still hands-on: you have recently designed and implemented production systems, not only managed people who didExperience working on horizontal multi-server architecture to deliver redundancy, scale, reliability and performanceProduction experience in Linux, Oracle and Java tech stackHands-on AWS core services (EC2, VPC, IAM, S3, RDS, CloudWatch): you can design accounts, networks, and IAM, and reason about cost, scaling, and resilience.Solid hypervisor/VM concepts (KVM, VMware, or similar); Proxmox familiarity a plusYou have implemented (not only selected) monitoring/observability and CI/CDHands-on access control/IAM, encryption, patching, vulnerability management; HIPAA, SOC 2, HITRUST, or ISO 27001 strongly preferredExperience leading or closely partnering with production teams based