Full Time
Our client is seeking a Senior Cloud Engineer to support and optimize large-scale cloud-based production environments, with a strong focus on reliability, operational excellence, and incident management.
This role is ideal for engineers with deep hands-on experience managing cloud infrastructure in production, whether in AWS, GCP, or Azure environments. You will be responsible for troubleshooting complex platform issues, supporting Kubernetes workloads, automating operational processes, and helping maintain highly available systems that support critical business applications.
What You'll Be Doing
- Investigate, troubleshoot, and resolve complex infrastructure and platform issues across cloud, networking, compute, storage, and application layers.
- Lead root cause analysis efforts for critical production incidents and implement preventative measures to reduce recurrence.
- Manage and support cloud-based infrastructure and containerized workloads running in production environments.
- Maintain, optimize, and troubleshoot Kubernetes clusters and cloud-native services.
- Build and improve automation solutions to reduce manual operational tasks and improve system reliability.
- Provision and manage infrastructure using Infrastructure as Code (IaC) tools and modern cloud engineering practices.
- Serve as a senior technical escalation point for complex operational and platform-related issues.
- Review infrastructure changes and provide guidance on operational best practices and system reliability.
- Collaborate with engineering, platform, and operations teams to coordinate incident response, change management, and service improvements.
- Help prioritize operational workstreams and contribute to maintaining stable, resilient production environments.
- Participate in an on-call rotation supporting business-critical systems and services.
What We're Looking For
- Strong experience troubleshooting and resolving production issues across cloud infrastructure, networking, operating systems, containers, and distributed systems.
- Hands-on experience operating cloud environments in AWS, GCP, or Azure, with a solid understanding of cloud architecture, networking, security, and operational best practices.
- Proven experience managing Kubernetes-based workloads in production environments.
- Experience provisioning and maintaining infrastructure using Infrastructure as Code tools such as Terraform or similar platforms.
- Strong understanding of incident management, post-incident reviews, root cause analysis, and preventative action planning.
- Ability to work effectively in high-pressure situations and drive resolution of critical service-impacting incidents.
- Strong communication and collaboration skills, with the ability to work across technical teams and stakeholders.
- Experience mentoring engineers or providing technical guidance within an operations, platform, or cloud engineering function.
Desirable Experience
- Exposure to service mesh technologies and Kubernetes policy management tools.
- Experience implementing Kubernetes scaling and cluster optimization solutions.
- Scripting or software development experience using Python, Bash, Go, or similar languages.
- Cloud, Kubernetes, or infrastructure-related certifications.
- Experience supporting large-scale, highly available, cloud-native production systems.
Work Setup
- Full-time, Onsite
-
Rotational shift schedule with monthly rotation across morning, afternoon, and night shifts
- Hybrid work flexibility available on select weekends and holidays, subject to company policies
- Additional remote work privileges may be granted based on performance and internal eligibility requirements
- Participation in an on-call support rotation may be required depending on operational needs