Role Overview We are seeking an Operational Excellence Manager to drive standardization, continuous improvement, and governance across data center operations. This role is responsible for building and sustaining a high-performance operating model that ensures consistency, compliance, and reliability across all sites. The role partners closely with Site Leaders, Facilities, and Data Center Operations teams to embed best practices, improve processes, and enhance overall operational maturity in line with hyperscaler expectations.
Key Responsibilities
1. Operational Standards & Governance
- Develop, implement, and maintain standardized operating procedures (SOPs, MOPs, and EOPs).
- Ensure consistency in operations across multiple sites and regions.
- Establish governance frameworks for change management, incident management, and maintenance practices.
- Drive adherence to internal standards and customer requirements.
2. Continuous Improvement & Process Optimization
- Identify inefficiencies and drive process improvements across operations and facilities.
- Lead Lean Six Sigma initiatives to improve uptime, reduce errors, and increase efficiency.
- Standardize workflows for maintenance, incident response, and deployments.
- Track and drive closure of improvement actions resulting from audits, incidents, and operational reviews.
3. Performance Management & Metrics
- Define and track operational KPIs, including uptime, MTTR, SLA compliance, and maintenance performance.
- Build dashboards and reporting frameworks to provide visibility for site management and leadership teams.
- Analyze trends and identify systemic issues impacting performance.
- Drive data-driven decision-making across operations teams.
4. Audit, Compliance & Risk Management
- Lead internal and external audits, including ISO audits, customer audits, and regulatory inspections.
- Ensure audit readiness and compliance across all operational processes.
- Track and close audit findings and non-conformities.
- Drive risk identification, mitigation strategies, and control frameworks.
5. Incident & Problem Management Excellence
- Establish and enforce best practices for incident response and escalation management.
- Lead post-incident reviews (RCA) and ensure corrective and preventive actions are implemented.
- Identify recurring issues and drive permanent corrective solutions.
- Improve response times and minimize operational impact.
6. Training & Capability Development
- Develop training programs for Operations and Facilities teams.
- Ensure teams are trained, certified, and aligned with operational standards.
- Drive knowledge sharing and adoption of best practices across sites.
- Support onboarding and capability development for new sites and team members.
7. Tooling, Automation & Digitalization
- Drive adoption and optimization of DCIM, CMMS, and workflow management tools.
- Identify opportunities for automation across operations and reporting processes.
- Improve data quality, asset tracking, and operational visibility.
- Partner with IT and Digital teams to enhance operational systems and capabilities.
8. Cross-Functional Collaboration
- Work closely with:
- Site Leaders
- Facilities Management teams
- Data Center Operations teams
- Engineering & Design teams
- Ensure alignment between design intent, operational practices, and customer expectations.
Qualifications
Required Qualifications
- Bachelor's degree in Engineering, Operations Management, or a related field.
- 8-12+ years of experience in Data Center Operations, Critical Facilities, or a similar environment.
- Strong understanding of data center operations and maintenance practices.
- Experience in process improvement, quality management, or operational excellence roles.
Preferred Qualifications
- Certifications such as:
- Lean Six Sigma (Green Belt or Black Belt)
- ITIL
- CDCP / CDCS / CDCE
- Experience in hyperscale or large colocation environments.
- Familiarity with ISO standards such as ISO 9001, ISO 27001, and ISO 45001.
- Experience with DCIM and CMMS tools.
Key Skills & Competencies
- Process-driven mindset with strong attention to detail.
- Data analysis and performance management.
- Strong stakeholder management and influencing skills.
- Problem-solving and root cause analysis.
- Ability to drive change across teams without direct authority.
- Strong documentation and communication skills.
Success Metrics (KPIs)
- Improvement in SLA adherence and uptime metrics.
- Reduction in incidents and recurring failures.
- Audit success rate and timely closure of findings.
- Standardization across sites, measured through process compliance scores.
- Improvement in MTTR and incident response times.
- Training completion rates and certification levels across teams.