Apply Edge Start your job search

Sr Site Reliability Engineer (SRE)

CareCone Group · Sydney, New South Wales, Australia

Apply & track with Apply Edge

Role: Sr Site Reliability Engineer (SRE) Experience: 8-10 years

Location: Sydney Key TechnologiesAWS | EKS | ECS | EC2 | Lambda | GitLab | Terraform | Ansible | Python | Bash | Linux | Kubernetes | Docker | PostgreSQL | Oracle | Vault | CloudWatch | OpenTelemetry | Grafana | Splunk | ELK | SLI | SLO | SLA | DevOps | Site Reliability EngineeringRole SummaryWe are seeking an experienced AWS DevOps & Site Reliability Engineer (SRE) to build, automate, operate, and continuously improve cloud platforms and mission-critical applications.

The role focuses on AWS cloud engineering, infrastructure automation, CI/CD, platform reliability, observability, incident management, resiliency, and operational excellence within a highly regulated enterprise environment.The successful candidate will drive automation, reliability, performance, scalability, and availability of cloud-native platforms while partnering with engineering teams to improve release quality, operational stability, and customer experience.Key ResponsibilitiesCloud & Platform EngineeringDesign, build, and maintain AWS cloud infrastructure and platform services.Develop and manage CI/CD pipelines using GitLab.Support cloud migration and platform modernization initiatives.Implement Infrastructure as Code (Terraform) and configuration management (Ansible).Automate deployments, environment provisioning, database refreshes, and operational processes.Manage secrets, certificates, access controls, and cloud security controls.Develop reusable infrastructure modules, deployment standards, and platform engineering patterns.Site Reliability EngineeringDefine and implement reliability engineering practices, operational standards, and platform blueprints.Establish and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs).Improve platform reliability, scalability, resilience, and fault tolerance for mission-critical applications.Drive proactive reliability improvements through automation and elimination of operational toil.Conduct capacity planning, performance tuning, and resource optimization activities.Support disaster recovery planning, backup validation, resilience testing, and business continuity initiatives.Participate in production readiness reviews and ensure operational requirements are embedded into solution designs.Observability & MonitoringDesign and implement monitoring, logging, alerting, and observability frameworks.Build and maintain dashboards, health checks, metrics, and operational reporting.Enhance end-to-end system visibility using CloudWatch, Grafana, Splunk, ELK, OpenTelemetry, or similar technologies.Drive alert tuning and noise reduction to improve operational effectiveness.Establish monitoring standards across applications, infrastructure, databases, and integration components.Incident & Operational ManagementLead incident triage, troubleshooting, root cause analysis, and post-incident reviews.Support production systems and provide timely resolution of infrastructure and application issues.Drive reduction of Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR).Develop operational runbooks, knowledge articles, and recovery procedures.Collaborate with engineering and support teams to implement preventative actions and continuous improvement initiatives.Participate in on-call and major incident management activities where required.DevOps & Engineering EnablementPromote DevOps, SRE, and Platform Engineering best practices.Collaborate with development teams to improve deployment reliability and automation maturity.Integrate security, observability, and operational controls into CI/CD pipelines.Support release management, change governance, and deployment strategies.Champion Infrastructure as Code, automation, and self-service platform capabilities.Mandatory SkillsAWS CloudAWS EC2, ECS/EKS, Lambda, VPC, IAM, S3, CloudWatchRDS/AuroraRoute53Secrets ManagerAWS Networking and Security ServicesDevOps & AutomationGitLab CI/CDTerraformAnsibleBash and Python scriptingLinux/Unix AdministrationInfrastructure AutomationSite Reliability EngineeringProduction Support and Incident ManagementSLI / SLO / SLA implementation and managementReliability Engineering practicesHigh Availability and Fault-Tolerant Architecture DesignCapacity Planning and Performance OptimizationDisaster Recovery and Business ContinuityRoot Cause Analysis and Problem ManagementObservabilityMonitoring, Logging and AlertingCloudWatchGrafanaSplunk / ELKOpenTelemetryOperational Metrics and DashboardingPlatform & ContainersKubernetes (EKS preferred)Containerisation (Docker)Networking and Infrastructure TroubleshootingSecuritySecrets Management Solutions (Vault preferred)IAM and Access ControlsCloud Security Best PracticesPreferred SkillsAWS Certified Solutions Architect / DevOps Engineer CertificationCertified Kubernetes Administrator (CKA)Experience with Aurora PostgreSQL and Oracle migration programsChaos Engineering and Resilience TestingExperience implementing SRE operating modelsService Mesh technologiesFinancial Services, Banking, Capital Markets, or other regulated industriesITIL processes including Incident, Problem, Change and Release ManagementExperience supporting large-scale cloud-native microservices platforms