What we do:
This role involves using your expertise in infrastructure, automation, and observability to proactively maintain and enhance platform stability. You will collaborate closely with development and engineering teams to deliver high-quality and resilient software solutions.
Key Responsibilities :
- Design and implement build, deployment, and configuration management
- Improve the frameworks, techniques, and tooling to provide increased quality and agility for our Development teams
- Manage life cycle management of the private and public cloud or containers and VMs server/infrastructure and related platforms such as AWS and Tencent.
- Build and test automation tools for infrastructure provisioning, maintain, and monitor configuration standards
- Gather and analyze metrics from both operating systems and applications to assist in performance tuning and fault-finding.
- Assist with the design, implementation, and administration of shared development, monitoring, CI/CD, DevOps, and collaboration tools
- Partner with development teams to improve services through rigorous testing and release procedures.
- Participate in system design consulting, platform management, and capacity planning
- Create sustainable systems and services through automation and uplift.
- Collaborate with other engineers in order to troubleshoot various environments
- Teach and mentor junior Site Reliability Engineers
- Balance feature development speed and reliability with well-defined service level objectives.
Required Qualifications (What we need) :
- At least 3-4 years of experience in a Production availability or SRE role.
- At least 2-3 years of experience in a scripting language (such as Powershell, Bash, or Python)
- Experience in Monitoring system and design reliability platform
- Experience with modern languages development practices and frameworks (such as JavaScript/Java/NodeJS/Golang)
- Ability to understand and implement CI/CD pipelines on AWS
- Experience in containerizing and orchestrating Docker (preferably with Kubernetes, ECS, or equivalent)
- Experience with writing Infrastructure as Code using Terraform or Cloudformation
- Identify opportunities to reduce time to delivery, rework, and total cost of ownership while improving the functional and non-functional requirements of our system
- Previous experience building highly scalable SaaS or PaaS solutions in a production environment
- Worked in an agile environment and understands agile practices
- Experience working in a heavily regulated industry or environment