
We are Hiring for one of the Top Most Software Product Development Company in Hyderabad.
Please find the details as below -
Role: VP - Technology Operations
Service Reliability & Availability:
- Deliver Five Nines Availability (99.999%): Build and enforce architectural and operational practices that ensure global transit payment systems achieve and sustain ultra-high uptime. This includes proactive monitoring, high-availability design enforcement, and automated failover systems across cloud and on-premises platforms.
- Govern Service Level Objectives (SLOs) & Error Budgets: Define, track, and report SLOs for all critical services, ensuring error budgets are managed responsibly to balance reliability with change velocity.
Observability, Automation & SRE Practices
- Full-Stack Observability: Deploy and govern monitoring solutions across metrics, logs, and traces, enabling real-time visibility into system health and early anomaly detection.
- Automation of Toil: Drive a culture of automation by eliminating repetitive manual tasks in patching, scaling, failover, and incident response. Track progress with explicit automation coverage targets.
- Chaos Engineering & Resilience Testing: Institutionalize chaos testing and DR drills ("game days") to validate RTO/RPO readiness and system recovery under stress conditions.
Global Operations & Customer Support
- Global Ops Center Leadership: Direct two 24x7 global operations centers (UK and India) as the backbone of global service delivery. Ensure staffing, shift rotations, and runbooks meet the highest standards of responsiveness and reliability.
- Localized Team Oversight: Manage in-country support teams that ensure local compliance, and customer-specific responsiveness.
- Customer Contact Centers: Own the operations of customer-facing contact centers, ensuring tight integration with back-end service teams to provide consistent, rapid, and high-quality customer experience.
Didn’t find the job appropriate? Report this Job