Site Reliability Engineer - HFT
Role Overview
Key Responsibilities
- Own and enhance end-to-end observability across production systems.
- Design actionable monitoring dashboards, alerts, and operational metrics.
- Investigate production incidents, diagnose root causes, and implement permanent fixes.
- Manage deployments through CI/CD pipelines and GitOps workflows.
- Operate and optimize Kubernetes-based workloads in cloud environments.
- Provision and maintain infrastructure using Infrastructure-as-Code practices.
- Build and support near real-time data pipelines and telemetry platforms.
- Design data quality controls and optimize analytical databases for large-scale querying.
- Develop and improve SQL-based analytics and reporting capabilities.
- Participate in a shared on-call rotation supporting critical production services.
Key Requirements
- Degree in Computer Science, Software Engineering, or equivalent practical experience.
- Experience building and maintaining observability and monitoring solutions within production environments.
- Strong understanding of incident management, troubleshooting, and root cause analysis.
- Hands-on experience with monitoring tools such as Prometheus, Grafana, ELK, Datadog, CloudWatch, or similar.
- Experience deploying software through CI/CD platforms such as GitLab CI, Jenkins, or GitHub Actions.
- Exposure to distributed stream processing technologies such as Flink, Spark Structured Streaming, or comparable frameworks.
- Experience working with OLAP or columnar databases such as ClickHouse, BigQuery, Druid, or Redshift.
- Practical Kubernetes experience, including deployments, resource management, and application configuration.
- Strong programming skills in Python, Java, Scala, Rust, or similar languages.
- Advanced SQL skills with experience building and optimising analytical workloads.
- Knowledge of messaging platforms and event-driven architectures such as Kafka or SQS
Nice to Have
- Experience with OpenTelemetry and distributed tracing.
- Knowledge of Airflow or workflow orchestration platforms.
- Advanced Terraform and Infrastructure-as-Code expertise.
- Familiarity with JVM build tooling or large-scale software delivery pipelines.
- Prior exposure to trading, market data, financial technology, or other low-latency environments.
FAQs
Congratulations, we understand that taking the time to apply is a big step. When you apply, your details go directly to the consultant who is sourcing talent. Due to demand, we may not get back to all applicants that have applied. However, we always keep your resume and details on file so when we see similar roles or see skillsets that drive growth in organizations, we will always reach out to discuss opportunities.
Yes. Even if this role isn’t a perfect match, applying allows us to understand your expertise and ambitions, ensuring you're on our radar for the right opportunity when it arises.
We also work in several ways, firstly we advertise our roles available on our site, however, often due to confidentiality we may not post all. We also work with clients who are more focused on skills and understanding what is required to future-proof their business.
That's why we recommend registering your resume so you can be considered for roles that have yet to be created.
Yes, we help with resume and interview preparation. From customized support on how to optimize your resume to interview preparation and compensation negotiations, we advocate for you throughout your next career move.
