Senior SRE


Austin
Permanent
Negotiable
Investment Management
PR/604590_1785939847
Senior SRE

Senior Site Reliability Engineer

A leading investment management firm is seeking a Senior Site Reliability Engineer to help scale and support critical workflow orchestration and automation platforms across the organization. This role sits within a Platform Engineering team responsible for delivering highly available, resilient, and scalable infrastructure that powers business-critical workloads and data processes.

The ideal candidate will have a strong background in SRE, DevOps, or Platform Engineering and enjoy balancing hands-on operational support with long-term engineering improvements. You'll work closely with engineering, data, infrastructure, and security teams to enhance platform reliability, automate manual processes, and drive modernization efforts across the environment.

Responsibilities

  • Provide operational ownership of workflow orchestration and enterprise scheduling platforms, including Apache Airflow and similar technologies.
  • Act as an escalation point for platform-related incidents, troubleshooting complex production issues and driving root cause analysis through resolution.
  • Partner with application and data teams to resolve workflow failures, dependency issues, scheduling conflicts, and performance bottlenecks.
  • Develop and maintain reliability standards, service objectives, monitoring strategies, and operational best practices.
  • Build automation and tooling that reduce manual effort and improve the overall user experience for engineering teams.
  • Design and implement observability solutions utilizing metrics, dashboards, alerting, logging, and performance monitoring.
  • Support platform lifecycle management, including upgrades, patching, configuration management, and security remediation.
  • Contribute to infrastructure modernization initiatives involving cloud services, containerization, platform migrations, and deployment automation.
  • Develop and maintain Infrastructure-as-Code solutions using tools such as Terraform, Helm, Ansible, and related technologies.
  • Participate in architectural discussions, platform roadmap planning, and engineering standards development.
  • Maintain operational documentation, procedures, and on-call readiness for supported environments.

Required Experience

  • Bachelor's degree in Computer Science, Engineering, Information Systems, or equivalent experience.
  • 5+ years of experience in Site Reliability Engineering, Platform Engineering, Infrastructure Engineering, or DevOps-focused roles.
  • Strong production experience supporting Apache Airflow environments.
  • Experience with distributed Airflow deployments, including Celery and/or Kubernetes executors.
  • Experience supporting enterprise workload automation and job scheduling platforms such as Automic/UC4, Control-M, or comparable technologies.
  • Strong Linux administration skills with working knowledge of Windows-based environments.
  • Experience supporting cloud infrastructure, preferably within AWS environments.
  • Proficiency with Python and scripting for automation, tooling, and operational efficiency.
  • Experience working with Kubernetes, Docker, CI/CD pipelines, and modern deployment methodologies.
  • Strong understanding of monitoring, logging, tracing, and observability concepts.
  • Experience with tools such as Grafana, Prometheus, ELK, or comparable monitoring platforms.
  • Proven ability to manage production incidents and communicate effectively during high-priority situations.
  • Strong automation mindset with a focus on improving efficiency and reducing operational overhead.

Preferred Qualifications

  • Hands-on experience administering Broadcom Automic/UC4.
  • Experience with managed Airflow platforms such as AWS MWAA, Cloud Composer, or Astronomer.
  • Exposure to modern data ecosystems, including technologies such as Kafka, dbt, and Snowflake.
  • Experience migrating workloads from legacy scheduling platforms to cloud-native orchestration solutions.
  • Familiarity with SLOs, SLIs, error budgets, and reliability engineering best practices.

FAQs

Congratulations, we understand that taking the time to apply is a big step. When you apply, your details go directly to the consultant who is sourcing talent. Due to demand, we may not get back to all applicants that have applied. However, we always keep your resume and details on file so when we see similar roles or see skillsets that drive growth in organizations, we will always reach out to discuss opportunities.

Yes. Even if this role isn’t a perfect match, applying allows us to understand your expertise and ambitions, ensuring you're on our radar for the right opportunity when it arises.

We also work in several ways, firstly we advertise our roles available on our site, however, often due to confidentiality we may not post all. We also work with clients who are more focused on skills and understanding what is required to future-proof their business. 

That's why we recommend registering your resume so you can be considered for roles that have yet to be created. 

Yes, we help with resume and interview preparation. From customized support on how to optimize your resume to interview preparation and compensation negotiations, we advocate for you throughout your next career move.