About the Company
Role - Cloud Reliability Engineer
Location - PAN India
About the Role
Job Requirements
Responsibilities
- Define and maintain service SLIs, SLOs, error budgets, reliability dashboards, and operational scorecards.
- Design, build, and operate secure, highly available, scalable, and cost-conscious cloud platforms and services.
- Implement observability covering metrics, logs, traces, synthetic checks, business signals, alerting, and diagnostic context.
- Improve alert quality, reduce noise, and establish actionable escalation and incident-response processes.
- Lead or support incident response, root-cause analysis, corrective actions, and blameless postmortems.
- Automate provisioning, deployment, configuration, diagnostics, remediation, maintenance, and recovery activities.
- Implement self-healing, automated rollback, progressive delivery, canary, blue-green, and safe deployment practices.
- Operate and optimize Kubernetes workloads, autoscaling, workload placement, resource limits, health checks, and availability controls.
- Perform load, stress, soak, failover, chaos, and disaster-recovery testing to identify weaknesses before production impact.
- Partner with development teams to embed reliability, observability, security, and operability into the software lifecycle.
- Support cloud migration readiness, dependency assessment, baseline measurement, and post-migration reliability validation.
- Conduct capacity planning, performance tuning, resource optimization, and FinOps-aligned cost improvement.
- Maintain runbooks, architecture and support documentation, audit evidence, operational standards, and knowledge assets.
- Mentor engineers and promote SRE, automation, continuous learning, and shared operational ownership.
Qualifications
- 6-12 years of experience in Cloud Engineering, Site Reliability Engineering, DevOps, Platform Engineering, or Production Engineering.
- Hands-on experience operating production workloads on AWS, Azure, and/or GCP.
- Strong knowledge of SLIs, SLOs, error budgets, availability, latency, throughput, saturation, and reliability risk.
- Experience with Prometheus, Grafana, Dynatrace, CloudWatch, Azure Monitor, Google Cloud Monitoring, OpenTelemetry, ELK, Datadog, or comparable tools.
- Hands-on experience with Docker, Kubernetes, and managed services such as EKS, AKS, or GKE.
- Experience with Terraform or CloudFormation and reproducible environment provisioning.
- Strong CI/CD experience using Jenkins, GitHub Actions, GitLab CI, Azure DevOps, or equivalent tools.
- Proficiency in Python, Bash, Shell, PowerShell, or Go for automation, diagnostics, and operational tooling.
- Knowledge of incident response, root-cause analysis, blameless postmortems, problem management, and on-call practices.
- Experience with resilience patterns such as timeouts, retries, circuit breakers, bulkheads, graceful degradation, and automated rollback.
- Knowledge of high availability, disaster recovery, capacity planning, performance testing, chaos engineering, and failover validation.
- Understanding of IAM, least privilege, secrets, encryption, vulnerability remediation, auditability, and DevSecOps controls.
Required Skills
- Cloud migration and modernization reliability validation.
- Distributed microservices, API platforms, messaging, or high-volume transactional systems.
Preferred Skills
- Cloud migration and modernization reliability validation.
- Distributed microservices, API platforms, messaging, or high-volume transactional systems.