Apply Edge Start your job search

Senior Site Reliability Engineer - Cloud Platform | Energy Trading & Infrastructure Firm

Techfellow Limited · New York, United States

Apply & track with Apply Edge

[Up to c. $300k Comp Package | Hybrid Working - 4 Days In Office]Role OverviewWe’re representing a global commodities trading and investment organisation expanding the infrastructure capability behind its firmwide investment, trading and analytics technology platform. The environment combines cloud infrastructure, container platforms, internal applications, APIs and data-driven tooling used across business-critical workflows.This hire will sit within a small Platform Infrastructure function and take substantial hands-on ownership across reliability engineering, AWS architecture, Kubernetes, infrastructure automation and production observability. A major part of the mandate is moving the platform towards more resilient, repeatable cloud-native operating patterns while improving how services are deployed, monitored and recovered. Alongside core SRE responsibilities, the role offers meaningful ownership of the organisation’s OpenTelemetry adoption, Kubernetes evolution and disaster recovery capability. It suits an experienced engineer who wants to remain deeply technical while influencing how critical infrastructure is designed and operated across the firm...Role SnapshotBring 6+ years of hands-on experience across Site Reliability Engineering, Platform Engineering, DevOps or production infrastructure, with evidence of owning business-critical systems rather than primarily supporting themDesign and operate highly available AWS infrastructure, applying resilient patterns across compute, storage, networking, scaling, load balancing, backup and recoveryBuild and evolve Kubernetes-based platform infrastructure using strong practical experience with Kubernetes, Docker, Terraform and GitOwn Infrastructure as Code and delivery automation, creating repeatable deployment patterns and improving CI/CD, testing, upgrades, patching and environment consistencyDevelop the observability platform across metrics, logs and distributed tracing, including continued adoption of OpenTelemetry, alongside tooling such as DatadogEstablish measurable reliability practices covering SLOs, SLIs, error budgets, capacity, operational health and reduction of recurring production failureStrengthen incident management through effective alerting, troubleshooting, runbooks, root-cause analysis and engineering changes that prevent repeat incidentsDesign and validate high-availability and disaster recovery capabilities, including RTO/RPO targets, recovery procedures, failover testing and dependency planningUse Python, Go, Bash or comparable engineering automation to reduce manual infrastructure work, supported by strong Linux, networking and production troubleshooting fundamentals(Preferred) Experience with Argo CD, Helm, Karpenter, Crossplane, progressive delivery, serverless architectures, additional cloud platforms or security controls within regulated or financial-services environments...