
Search by job, company or skills
Overview
At Xtremax, we help government agencies and enterprises build robust, scalable, and future-ready digital systems. As an Operations Support Engineer (Public Sector), you will play a key role in designing, building, and maintaining critical cloud infrastructure platforms, ensuring seamless system performance, security, and high availability across public sector environments. This role combines hands-on cloud and virtualization management, Site Reliability Engineering (SRE) practices, automated infrastructure deployments, and proactive monitoring to ensure full compliance with Whole-of-Government(WOG) standards.
Candidates with public sector experience are preferred, as this role directly supports IT projects for government agencies.
Responsibility:
Infrastructure Management: Design, build, and maintain critical cloud infrastructure platforms encompassing compute, storage, networking, containerisation, virtualisation, DNS, monitoring, and supporting systems across development, staging, and production environments. Monitor and manage comprehensive cloud services including CloudWatch logs, alarms, synthetic monitoring, and integrated third-party solutions.
Monitoring and Observability: Implement and maintain robust monitoring and observability frameworks for all platform components utilising modern tooling including AWS CloudWatch Canaries, StackOps, Prometheus, Grafana, and ELK stack implementations. Establish comprehensive observability practices to support proactive problem diagnosis and provide actionable insights into system health and performance metrics.
Compliance and Security: Maintain adherence to Whole-of-Government platform standards, compliance frameworks, and security requirements through continuous monitoring using government-approved security and monitoring solutions. Implement security controls including access management, security hardening, and compliance monitoring with tools such as CyberArk.
Automation and Infrastructure as Code: Develop and maintain infrastructure using Infrastructure as Code (IaC) methodologies with tools including Terraform, Ansible, and AWS CloudFormation to ensure repeatable, automated, and version-controlled deployments. Follow platform standards whilst executing infrastructure automation and modern operational practices to enhance efficiency and reliability.
Site Reliability Engineering: Identify and eliminate repetitive operational tasks to improve Developer and Infrastructure Engineer efficiency whilst enhancing overall system reliability through systematic toil elimination and error budget management. Define, track, and report on SRE metrics including Service Level Objectives (SLO), Service Level Indicators (SLI), and error budgets.
Platform Operations: Manage virtualisation platforms including VMware vSphere and Hyper-V, encompassing capacity monitoring, performance optimisation, and lifecycle management. Administer AWS Cloud services including EC2, ECS, S3, RDS(PostgreSQL and MS SQL), Docker/Kubernetes, Lambda, CloudFormation, CloudWatch, IAM, and VPC configurations alongside physical server infrastructure.
Network and System Administration: Demonstrate proficiency with local networking technologies including TCP/IP, DNS, DHCP, VPN configurations, and routing protocols. Execute comprehensive platform patching strategies leveraging automation to maintain security and stability whilst minimising service disruption.
Business Continuity: Maintain backup, disaster recovery, and high availability solutions for critical platform components including AWS Fault Injection Simulator (FIS) testing and multi-availability zone configurations. Support containerisation initiatives and maintain container orchestration platforms for traditional workloads.
Collaboration and Documentation: Collaborate effectively with application teams to support platform stability, performance, and scalability requirements. Create and maintain comprehensive platform documentation, operational runbooks, and standard operating procedures. Support team development through knowledge sharing and
mentoring on platform operations and modern infrastructure practices.
Requirements:
Technical Expertise:
Professional Qualifications: Bachelor's degree in computer science, Information Technology,or related technical discipline with demonstrated experience in infrastructure operations and engineering. Strong understanding of enterprise infrastructure components with proven experience supporting infrastructure modernisation initiatives.
Core Competencies: Excellent analytical and problem-solving capabilities with strong documentation skills and effective communication abilities for both technical and non-technical stakeholders.
Desired Certifications:
Job ID: 151685153