🎁 Before you apply, rehearse this interview. Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card. I want my training →
Job Summary We are looking for a highly hands-on Head of Infrastructure & Cloud to lead and operate the company’s cloud, on-premise and production infrastructure environments. This is not a management-only role. The successful candidate must be capable of personally designing, implementing, troubleshooting and operating critical infrastructure while leading infrastructure personnel, vendors and operational activities. The role is responsible for ensuring the availability, performance, scalability, security and cost effectiveness of business-critical systems, with a target of at least 99.9% monthly service availability. The successful candidate will take ownership of production operations, including major incidents, service disruptions, capacity planning, production changes, backup, Business Continuity and Disaster Recovery. The role reports directly to the CTO and acts as the company’s principal infrastructure and cloud subject matter expert. This is a fully on-site role requiring attendance at our Kuala Lumpur office. Key Responsibilities Develop and execute the company’s cloud, infrastructure and production-operations strategy. Design, operate and continuously improve infrastructure across AWS, Huawei Cloud and on-premise environments. Personally troubleshoot and resolve complex production, cloud, network and infrastructure issues when required. Manage RedHat/AmazonLinux/EulerOS Linux systems, virtualisation, networking, load balancing, DNS, SSL/TLS, storage and application infrastructure. Manage Cloudflare services including DNS, CDN, WAF, DDoS protection, Waiting Room, Load Balancing and traffic management. Ensure critical systems achieve or exceed 99.9% monthly service availability. Establish and operate monitoring, alerting, observability, capacity management and operational dashboards. Lead P1 and other major production incidents, including troubleshooting, stakeholder coordination, recovery and post-incident reviews. Conduct capacity planning, infrastructure optimisation and performance tuning. Support JVM, application and MySQL performance analysis where infrastructure involvement is required. Develop and maintain Business Continuity and Disaster Recovery strategies, including RTO, RPO, backup, replication, failover and failback. Conduct regular DR tests, recovery exercises and tabletop exercises. Establish secure and controlled CI/CD, infrastructure and production-deployment processes. Implement Infrastructure as Code and automation using tools such as Terraform, Ansible, Bash or Python. Manage production access, privileged accounts, secrets, certificates and infrastructure security controls. Ensure production changes are properly reviewed, approved, auditable and reversible. Work closely with Cybersecurity and Engineering teams to protect production systems and data. Manage infrastructure-related cloud and vendor costs and identify opportunities to improve utilisation and reduce unnecessary expenditure. Maintain infrastructure architecture, documentation, operational procedures and recovery runbooks. Manage cloud, data-centre, network and infrastructure vendors. Provide the CTO with clear visibility of infrastructure risks, incidents, capacity issues, costs and unresolved operational problems. Required Skills and Experience Strong hands-on experience designing, operating and troubleshooting production infrastructure. Strong experience with AWS, Huawei Cloud or another major cloud platform. Strong Linux systems administration capability. Strong understanding of TCP/IP, DNS, routing, VPNs, firewalls, load balancing and network troubleshooting. Experience operating production systems with at least 99.9% SLA requirements. Strong monitoring, observability, incident-response and root-cause analysis experience. Experience with Business Continuity, Disaster Recovery, backup and recovery testing. Experience conducting or participating in DR exercises and tabletop exercises. Experience with Infrastructure as Code, preferably Terraform. Proficiency in scripting using Bash, Python or similar languages. Experience with CI/CD and production-deployment processes. Experience with infrastructure security, IAM, privileged access and secrets management. Strong capacity planning, performance optimisation and cost-management capability. Ability to operate independently and remain technically hands-on. Leadership and Professional Judgement Experience leading infrastructure, cloud, DevOps, systems or production-operations personnel. Ability to lead a small team while remaining hands-on. Set priorities, review technical work and ensure important operational actions are completed. Maintain close alignment with the CTO on infrastructure priorities, risks and major decisions. Challenge unsafe, unnecessarily expensive or technically weak proposals constructively. Balance availability, performance, security, cost and delivery when making decisions. Communicate infrastructure risks and incidents transparently. Maintain sound judgement and remain calm during high-pressure production incidents. Take ownership of outcomes and follow issues through to resolution. Availability and Accountability Participate in an on-call or emergency escalation arrangement. Be available during major production incidents, outages and infrastructure emergencies, including outside normal working hours when necessary. Remain contactable when the CTO or CEO requires infrastructure-related input, advice or support. Personally support resolution when production systems are materially affected. Support urgent customer complaints caused by infrastructure or production issues. Conduct post-incident reviews and ensure corrective actions are completed. Preferred Experience AWS and Huawei Cloud. Cloudflare Business or Enterprise. VMware. Red Hat Enterprise Linux / Amazon Linux. Nginx and Apache Tomcat. JVM and garbage-collection tuning. MySQL performance tuning. Terraform and Ansible. GitHub Actions. Prometheus, Grafana, PRTG or similar monitoring platforms. Hybrid cloud and on-premise infrastructure. High-volume or high-availability production systems. Relevant certifications such as AWS, Huawei Cloud, RHCSA/RHCE, VMware VCP, Terraform or networking certifications will be an advantage. Practical production capability, infrastructure troubleshooting, incident-management experience and demonstrated results will be given greater consideration than the number of certifications held. Candidate Profile We welcome strong mid-career or senior professionals who may not yet have held a Head of Infrastructure title, provided they demonstrate: Strong hands-on technical capability; Real production-operations experience; High-availability and DR experience; Strong troubleshooting and root-cause analysis; Sound technical judgement; Strong ownership and urgency; Cost awareness; Leadership capability; Calm decision-making under pressure; and A proven ability to keep critical systems reliable and available. For immediate consideration, kindly apply online.