HamburgerMenu
iimjobs

Posted by

Teja

Snr Hr at H2H Solutions

Last Active: 17 August 2026

Job Views:  
425
Applications:  226
Recruiter Actions:  1

Posted in

IT & Systems

Job Code

1724238

Head - Technology & Service Operations

premium_icon
H2H Solutions.15 - 30 yrs.Hyderabad
Posted 4 days ago
Posted 4 days ago

Position Summary

The Head of Technology and Service Operations is a senior executive responsible for the global operations, service delivery, and reliability of mission-critical transit payment systems used by millions of passengers daily. The role leads more than 500 staff across two global 24x7 operations centers (UK and India), localized in-country teams, and multiple customer contact centers.

Key Responsibilities

Service Reliability & Availability:

- Deliver Five Nines Availability (99.999%): Build and enforce architectural and operational practices that ensure global transit payment systems achieve and sustain ultra-high uptime. This includes proactive monitoring, high-availability design enforcement, and automated failover systems across cloud and on-premises platforms.

- Govern Service Level Objectives (SLOs) & Error Budgets: Define, track, and report SLOs for all critical services, ensuring error budgets are managed responsibly to balance reliability with change velocity.

- Transaction Performance Management: Guarantee low-latency, high-throughput processing across all environments, actively tuning Oracle databases, Kubernetes clusters, UCS fabrics and cloud services for peak demand conditions (e.g., major city events, commuter surges).

- Preventative Maintenance Program: Own a proactive, structured maintenance strategy (firmware updates, DB patching, load balancing, failover readiness) to ensure reliability and reduce risk of unplanned outages.

Incident, Problem & Change Management

- Global Incident Oversight: Lead the 24x7 incident response process across UK and India operations centers, ensuring - 95% SLA compliance for response and resolution.

- Root Cause & Blameless Postmortems: Mandate blameless postmortems for every significant incident, driving systemic fixes to prevent recurrence and sharing lessons learned across all regions.

- Problem Management: Establish a structured problem management process to identify trends, reduce repeat incidents, and address root technical or process flaws.

- Change Governance: Oversee change management to balance service stability with innovation. Implement automated pipelines where possible to reduce manual error and ensure changes are tested for resilience before release.

- Recovery Metrics: Continuously track and drive down Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR) with quarterly improvement goals.

Observability, Automation & SRE Practices

- Full-Stack Observability: Deploy and govern monitoring solutions across metrics, logs, and traces, enabling real-time visibility into system health and early anomaly detection.

- Automation of Toil: Drive a culture of automation by eliminating repetitive manual tasks in patching, scaling, failover, and incident response. Track progress with explicit automation coverage targets.

- Chaos Engineering & Resilience Testing: Institutionalize chaos testing and DR drills ("game days") to validate RTO/RPO readiness and system recovery under stress conditions.

- Capacity & Performance Engineering: Oversee predictive capacity planning to ensure the system scales automatically to meet load and maintain performance even under extreme demand conditions.

- Proactive Reliability Improvements: Fund and prioritize engineering initiatives aimed at improving long-term service resilience, not just reactive firefighting.

Global Operations & Customer Support

- Global Ops Center Leadership: Direct two 24x7 global operations centers (UK and India) as the backbone of global service delivery. Ensure staffing, shift rotations, and runbooks meet the highest standards of responsiveness and reliability.

- Localized Team Oversight: Manage in-country support teams that ensure local compliance, and customer-specific responsiveness.

- Customer Contact Centers: Own the operations of customer-facing contact centers, ensuring tight integration with back-end service teams to provide consistent, rapid, and high-quality customer experience.

- Standardization Across Regions: Drive consistency of ITIL-aligned service management processes globally, ensuring customers experience the same high standards regardless of geography.

Compliance, Risk & Audit Readiness

- Global Compliance Ownership: Ensure operations remain compliant with ISO27001, PCI DSS v4.1, Essential Eight, Cyber Essentials, GDPR, and all other local privacy regulations in customer jurisdictions.

- Audit Readiness & Evidence: Maintain continuous audit readiness with robust evidence trails across all systems, processes, and controls, avoiding last-minute remediation before audits.

- Corporate CISO & Service Assurance Partnership: Work in lockstep with the Corporate CISO and Service Assurance functions to embed security controls, risk frameworks, and governance processes into day-to-day operations, aiming for regulatory compliance excellence.

- Risk Reporting: Provide transparent reporting to the COO and Senior Leadership Teams on compliance posture, vulnerabilities, audit results, and remediation progress.

- Vulnerability Management: Ensure vulnerabilities are tracked, prioritized, and remediated on schedule, with operational accountability for timely fixes.

Financial Management & Vendor Oversight

- Budget Accountability: Own the operations budget, ensuring disciplined financial management, proactive forecasting, and cost optimization across global operations.

- Abatements & Liquidated Damages: Minimize exposure to penalties by delivering against contractual SLAs, proactively addressing risks that might result in service-level breaches.

- Vendor & Partner Management: Oversee the performance of critical vendors and partners - including telcos, cloud providers, and managed service providers. Hold them accountable to contractual SLAs and enforce remedies when necessary.

- Commercial Strategy: Negotiate favorable terms for renewals and expansions, aligning vendor performance with global service reliability goals.

- Cost Optimization: Continuously evaluate opportunities to optimize spending without compromising reliability or compliance.

Leadership & Governance

- Executive Leadership: Act as a key member of the Senior Leadership Team, reporting to the COO, with direct accountability for global technology and service operations.

- Cross-Functional Collaboration: Partner with the CTO and Software Delivery leaders to align infrastructure improvements, operational enhancements, and vulnerability remediation with technology roadmaps.

- Cadence of Accountability: Establish and enforce disciplined operational cadences (daily incident reviews, weekly ops reviews, monthly resilience drills, quarterly DR game days, annual audit cycles).

- Culture of Excellence: Foster a global culture of reliability-first thinking, preventative maintenance, process rigor, and continuous improvement.

- Talent Leadership: Build and develop global operational leadership talent, ensuring succession planning, skills development, and team engagement across all regions.

Candidate Profile

- Leadership: 15+ years in global technology operations, with 10+ years in senior leadership of large-scale, 24x7 organizations.

- Technical Breadth: Strong working knowledge of Azure, AWS, Kubernetes, Oracle databases, UCS fabrics, and enterprise networks; proven track record delivering highly available, resilient platforms.

- Service & Reliability Engineering: Deep experience with ITIL service management integrated with SRE practices (SLOs, error budgets, observability, automation, chaos testing).

- Compliance & Risk: Demonstrated success managing ISO27001, PCI DSS, GDPR, and local privacy obligations; experienced in working directly with CISOs and Service Assurance functions.

- Financial & Vendor Management: Proven ability to manage large budgets, minimize penalties, and hold global/cloud/telco vendors accountable to SLAs.

- Leadership Attributes: Process-driven operator with a reliability-first mindset; skilled communicator with executive presence; committed to building high-performing, globally distributed teams.

Didn’t find the job appropriate? Report this Job

Similar jobs that you might be interested in

Posted by

Teja

Snr Hr at H2H Solutions

Last Active: 17 August 2026

Job Views:  
425
Applications:  226
Recruiter Actions:  1

Posted in

IT & Systems

Job Code

1724238

Loading chat...