Senior Site Reliability Engineer

e-Hireo

JOB DESCRIPTION


Experience : 5 - 8 Yrs

Location : Bengaluru

Designation : Senior Site Reliability Engineer


Job Duties and Responsibilities:

  • Serve as the front line for infrastructure alerts and Incident Management in a global 24×7 follow-the-sun and On-Call model, ensuring rapid response, triage, escalation, service restoration, and effective cross-region handoffs.
  • Own and improve Customer Experience and Operational KPIs, including Customer Issue and DOHD ticket aging, response time, backlog burn-down, escalations, routing quality, go-live incidents, and RCA SLA compliance.
  • Drive daily triage of new, aging, blocked, and escalated work, ensuring clear ownership and timely closure.
  • Partner with Cloud Platform and Engineering to identify recurring customer issues and incident patterns, driving permanent fixes, automation, and preventive improvements.
  • Partner with Customer Enablement and cross-functional teams to establish clear ownership, collaboration channels, and points of contact for critical customer activities.
  • Improve the DOHD process for Customer Issues, CE questions, and requests by increasing ticket capture, reducing mis-routing, and improving response times.
  • Support critical customer go-lives, feature enablement, and per-tenant infrastructure requirements, ensuring operational readiness, risk management, and rollback planning.
  • Strengthen internal RCA governance through consistent DOHD workflows, SLA tracking, and timely closure of corrective and preventive actions.'


Skills You Must Have:

  • Engineering degree in Computer Science or a related technical field.
  • 5+ years of experience in SRE, SRE, Cloud Operations, Platform Engineering, or Production Engineering.
  • Strong experience supporting highly available SaaS or cloud platforms in a 24×7 production environment.
  • Hands-on experience with one or more major cloud providers: AWS, Google Cloud Platform, or Microsoft Azure.
  • Strong experience with Kubernetes and containerized production environments.
  • Experience with Infrastructure as Code, preferably Terraform, along with Jekins, GitOps, and automation practices.
  • Strong Linux, networking, troubleshooting, and distributed systems fundamentals.
  • Experience with observability, Incident Management, On-Call operations, RCA, SLA, SLO, and production reliability practices.
  • Strong analytical ability to define, manage, and improve operational KPIs and translate trends into measurable corrective actions.
  • Excellent communication and cross-functional collaboration skills, particularly during customer-critical situations and production incidents.
  • Strong ownership mindset with a focus on customer outcomes, operational excellence, and continuous improvement.

How to apply

To apply for this job you need to authorize on our website. If you don't have an account yet, please register.