Site Reliability Engineer


London
Permanent
Negotiable
Financial Technology
PR/605294_1787918930
Site Reliability Engineer

Our client, a world renowned hedge fund, is seeking a Site Reliability Engineer to join its world-class engineering organization in London. This role sits at the intersection of software engineering and infrastructure, focusing on the reliability, scalability, and performance of the technology platforms that power global trading and investment operations.

Working closely with software engineers, quantitative researchers, traders, and infrastructure teams, you will be responsible for building automation, improving observability, and ensuring critical production systems operate at the highest levels of availability and efficiency.

Key Responsibilities

  • Design, build, and maintain highly reliable, scalable, and automated infrastructure platforms.
  • Drive improvements in system performance, monitoring, observability, and operational efficiency.
  • Troubleshoot and resolve complex production incidents across distributed systems.
  • Develop tools and automation to reduce operational overhead and improve platform resilience.
  • Partner with engineering teams to improve system design, deployment processes, and operational readiness.
  • Participate in incident management and root cause analysis, ensuring lessons learned are incorporated into future improvements.
  • Support mission-critical trading and research environments in a fast-paced, high-performance setting.

Requirements

  • Strong software engineering skills in Python, Go, C++, Java, or a similar language.
  • Deep Linux systems knowledge and experience operating large-scale production environments.
  • Experience with Kubernetes, containerisation technologies, and cloud infrastructure.
  • Strong understanding of networking, distributed systems, and infrastructure automation.
  • Experience with monitoring and observability tools such as Prometheus, Grafana, Splunk, or similar.
  • Proven track record of solving complex reliability, scalability, or performance challenges.
  • Excellent problem-solving skills and ability to operate effectively in high-pressure environments.

Preferred Backgrounds

  • Technology companies operating large-scale distributed systems.
  • High-frequency trading firms, hedge funds, or electronic trading environments.
  • Cloud infrastructure, platform engineering, or production engineering teams.
  • Software engineers with a strong interest in reliability and infrastructure.

FAQs

Congratulations, we understand that taking the time to apply is a big step. When you apply, your details go directly to the consultant who is sourcing talent. Due to demand, we may not get back to all applicants that have applied. However, we always keep your resume and details on file so when we see similar roles or see skillsets that drive growth in organizations, we will always reach out to discuss opportunities.

Yes. Even if this role isn’t a perfect match, applying allows us to understand your expertise and ambitions, ensuring you're on our radar for the right opportunity when it arises.

We also work in several ways, firstly we advertise our roles available on our site, however, often due to confidentiality we may not post all. We also work with clients who are more focused on skills and understanding what is required to future-proof their business. 

That's why we recommend registering your resume so you can be considered for roles that have yet to be created. 

Yes, we help with resume and interview preparation. From customized support on how to optimize your resume to interview preparation and compensation negotiations, we advocate for you throughout your next career move.

Handpicked roles for you