Technical Skills Requirement : Linux | Shell Scripting | Git | CI/CD | Docker | Kubernetes |Observability (Logging, Metrics, Tracing & Alerting) | SIEM | Incident Response | APIs | Microservices | Secrets Management | Production Reliability(SLIs/SLOs) | Enterprise Platform Operations | Infrastructure & Cloud Troubleshooting
Experience: 5-8 years
Responsibilities
- Deploy and operationalize enterprise platform solutions across Dev, QA, UAT, and Production environments.
- Establish production-grade reliability for platforms, including logging, observability, monitoring, alerting, failure handling, and operational resilience.
- Integrate platforms with enterprise SIEM and Incident Response workflows for monitoring, alerting, triage, and escalation.
- Implement environment management, configuration management, release processes, and deployment automation.
- Ensure platforms meet enterprise standards for availability, scalability, security, recoverability, and operational support.
- Support production readiness reviews, pilot operations, stability testing, and handover to support teams.
- Troubleshoot platform, integration, networking, performance, and reliability issues across environments.
Technical Skills Required
- 4+ years of hands-on experience as an SRE, DevOps Engineer, Platform Engineer, Infrastructure Engineer, or Production Engineer.
- Hands-on experience deploying and operating production systems across multiple environments such as Dev, QA, UAT, and Prod.
- Strong knowledge of Linux, shell scripting, Git, CI/CD, Docker, and Kubernetes.
- Experience with observability platforms, logging, metrics, tracing, alerting, and dashboarding.
- Experience integrating systems with enterprise monitoring, alerting, SIEM, or Incident Response workflows.
- Experience defining and implementing runbooks, operational procedures, escalation paths, and production support models.
- Strong troubleshooting skills across application, infrastructure, networking, container, and cloud layers.
- Familiarity with production reliability practices such as SLIs, SLOs, incident management, capacity planning, failure handling, and post-incident review.
- Experience working with APIs, gateways, microservices, service accounts, secrets management, and enterprise integration patterns.
- Ability to collaborate with cybersecurity, infrastructure, application development, and operations teams.
Desired Skills
- Experience with automated testing, regression testing, security testing, or chaos testing.
- Experience with cloud platforms such as AWS, Azure, or Google Cloud.
- Experience with infrastructure-as-code tools such as Terraform, Helm, Argo CD, or similar deployment automation platforms.
- Familiarity with identity and access management, RBAC, OAuth/OIDC, secrets management, certificates, and network security.
- Experience operating regulated, high-availability, or enterprise-scale platforms.
- Working knowledge of a scripting/programming language (e.g., Python or Go) for automation and tooling.