Apply Edge Start your job search

Senior Site Reliability Engineer - HPC / On-premise Kubernetes

Opus Recruitment Solutions · Greater London, England, United Kingdom

Apply & track with Apply Edge

GPU | HPC | High Performance Computing | Data Centres | AI | DevOps | Observability | Site Reliability Engineer | Kubernetes | Staff & Principal Observability Platform EngineerExcited by the scaling of AI through Tech? Want to be part of that journey?Where better than someone building one of the largest GPU and HPC environments on the planet?Trusted by the likes of OpenAI they've raised over £bn in funding and are backed by some of the biggest names in Tech globally.With this incredible growth now couldn't be a more exciting time to get on board as they hire some integral roles across their Observability function. Working on all things Opensource they're service over tech and are hosting tools like Kubernetes, VictoriaMetrics, Loki, Prometheis, Grafana, Thanos, OpenTelemetry, Terraform IaC and more. Not just part of those teams but building and owning these technologies and function!You'll have come from somewhere that really can't tolerate downtime and somewhere that's heavy heavy heavy on metrics, logs, traces and alerting for an incredibly large-scale Kubernetes workload.If you're passionate about open-source technologies, large-scale systems and solving problems that don't have off-the-shelf answers, this is well worth a conversation.Apply now or contact robin.shaw@opusrs.comGPU | HPC | High Performance Computing | Data Centres | AI | DevOps | Observability | Site Reliability Engineer | Kubernetes |