Insurance
Help shape the reliability and scalability of a modern cloud-native platform used by thousands of users every day. As a Site Reliability Engineer, you’ll work alongside experienced Platform, DevOps and Software Engineers to automate infrastructure, improve system resilience and ensure high availability across production environments.
What You’ll Be Doing
- Build, maintain and optimise cloud infrastructure
- Manage and improve Kubernetes environments
- Develop Infrastructure as Code using Terraform
- Automate operational tasks with Ansible
- Monitor platform health and ensure high availability
- Troubleshoot production incidents and perform root cause analysis
- Improve CI/CD pipelines and deployment processes
- Work closely with Development, Security and Platform teams
- Drive reliability, scalability and automation initiatives
Tech Stack
- Kubernetes
- Terraform
- Ansible
- AWS
- Docker
- Helm
- Linux
- Prometheus
- Grafana
- ELK Stack
- ArgoCD
- GitHub Actions
- Git
- Bash
- Python
- CI/CD
- Infrastructure as Code
What We’re Looking For
- 3+ years of experience in Site Reliability Engineering, DevOps or Platform Engineering
- Strong hands-on experience with Kubernetes
- Experience managing cloud infrastructure on AWS
- Solid Terraform and Infrastructure as Code experience
- Experience with Ansible automation
- Good Linux administration skills
- Experience with monitoring and observability tools
- Familiarity with CI/CD pipelines and Git workflows
- Strong troubleshooting and problem-solving skills
- Excellent communication and collaboration skills