Cloud / Platform Operations Manager (AWS + Kubernetes)
Cloud / Platform Operations Manager (AWS & Kubernetes)
About the Role
We are seeking a hands-on Cloud / Platform Operations Manager to lead the day-to-day operations of our AWS and Kubernetes environments. This role is focused on platform reliability, cloud governance, service delivery, and operational excellence rather than software development. The ideal candidate will be a strong technical leader who thrives in a high-volume support environment, can resolve complex production issues, and is passionate about building and developing high-performing teams.
Key Responsibilities
- Lead daily operations, performance, and availability of AWS and Kubernetes platforms.
- Provide hands-on support and technical leadership for production infrastructure issues.
- Manage a ticket-driven support organization, ensuring timely response and resolution of user requests.
- Establish and maintain operational SLAs/SLOs, monitoring, alerting, and incident response processes.
- Own AWS governance, including IAM, access controls, security guardrails, compliance, and cost management.
- Maintain and optimize Kubernetes environments, including upgrades, patching, troubleshooting, networking, storage, and performance tuning.
- Serve as the primary escalation point for critical incidents and lead root cause analysis and remediation efforts.
- Lead, mentor, and develop a team of cloud and platform engineers while fostering a culture of accountability and customer service.
- Partner with engineering, product, security, and business teams to support platform needs and drive continuous improvement.
Required Qualifications
- 7+ years of cloud infrastructure, platform operations, or SRE experience, including 2+ years of team leadership.
- Strong hands-on experience supporting and administering production Kubernetes environments.
- Deep knowledge of AWS, including IAM, networking, governance controls, and multi-account environments.
- Experience implementing cloud security guardrails, policies, and operational best practices.
- Expertise with monitoring and observability tools such as Prometheus, Grafana, CloudWatch, or similar.
- Proven success leading operational support teams in high-volume, production environments.
- Strong incident management, troubleshooting, and stakeholder communication skills.
- Ability to remain technically hands-on while leading and developing a team.
Preferred Certifications: AWS Certifications, Kubernetes Certifications (CKA, CKAD, or CKS).
This is an excellent opportunity to lead a mission-critical cloud operations function supporting highly available Kubernetes and AWS platforms in a fast-paced, enterprise environment.
FAQs
Congratulations, we understand that taking the time to apply is a big step. When you apply, your details go directly to the consultant who is sourcing talent. Due to demand, we may not get back to all applicants that have applied. However, we always keep your resume and details on file so when we see similar roles or see skillsets that drive growth in organizations, we will always reach out to discuss opportunities.
Yes. Even if this role isn’t a perfect match, applying allows us to understand your expertise and ambitions, ensuring you're on our radar for the right opportunity when it arises.
We also work in several ways, firstly we advertise our roles available on our site, however, often due to confidentiality we may not post all. We also work with clients who are more focused on skills and understanding what is required to future-proof their business.
That's why we recommend registering your resume so you can be considered for roles that have yet to be created.
Yes, we help with resume and interview preparation. From customized support on how to optimize your resume to interview preparation and compensation negotiations, we advocate for you throughout your next career move.
