
Role & Responsibilities:
Cloud & Data & AI/ML Platform Operations:
- Own the operational health, availability, and performance of enterprise cloud platforms (AWS, Azure, GCP) ensuring production-grade reliability for models, pipelines, data & analytics platforms, and cloud-native applications.
- Establish and enforce MLOps and AIOps operational standards including model monitoring, drift detection, automated retraining pipelines, inference infrastructure management, and incident response for AI/ML workloads.
Site Reliability Engineering (SRE):
- Build, lead, and mature the enterprise SRE function embedding reliability engineering principles (SLOs, SLIs, error budgets, chaos engineering) across critical digital platforms and services.
- Lead the post-incident review (PIR) and blameless retrospective culture, ensuring every significant incident drives lasting systemic improvements rather than short-term fixes.
Managed Service Provider (MSP) Governance:
- Serve as the executive owner of all MSP and third-party operational vendor relationships governing contracts, SLAs, performance metrics, and strategic alignment across managed infrastructure, cloud, support, and security services.
- Lead structured QBRs, performance reviews, and executive-level escalations with MSP partners holding providers accountable to contractual commitments while fostering collaborative, long-term partnerships.
Enterprise IT Support & Service Management:
- Oversee the enterprise IT support function including Tier 1/2/3 support, service desk operations, and application support ensuring exceptional end-user experience and first-contact resolution metrics.
- Lead the continuous maturation of ITSM processes (Incident, Problem, Change, Release, and Configuration Management) in alignment with ITIL best practices and enterprise risk controls.
Technical Competencies:
- Deep expertise in cloud platform operations (AWS, Azure, GCP) architecture patterns, operational tooling, FinOps, and multi-cloud governance.
- Strong grounding in AI/ML operations including MLOps pipelines, model monitoring, inference infrastructure, and AIOps platform tooling.
- SRE mastery including observability stacks (Datadog, Dynatrace, Prometheus/Grafana), chaos engineering, SLO frameworks, and incident management platforms.
- ITSM fluency including ServiceNow or equivalent, ITIL v4 processes, and enterprise support operations at scale.
- MSP governance and vendor management including SLA construction, performance metrics, contract lifecycle, and strategic sourcing principles.
- Security and compliance operations awareness including vulnerability management, cloud security posture, and regulatory compliance frameworks.
Didn’t find the job appropriate? Report this Job