Principal SRE Engineer
Entain
Company Description
Entain India is the engineering and delivery powerhouse for Entain, one of the world's leading global sports and gaming groups. Established in Hyderabad in 2001, we've grown from a small tech hub into a dynamic force, delivering cutting-edge software solutions and support services that power billions of transactions for millions of users worldwide.
Our focus on quality at scale drives us to create innovative technology that supports Entain's mission to lead the change in global sports and gaming sector. At Entain India, we make the impossible possible, together.
Job Description
We are seeking a Principal Site Reliability Engineer(SRE) to set the technical direction for the reliability, performance, and scalability of our hybrid infrastructure. As the most senior technical authority in the SRE function, you will define architecture and engineering standards across cloud and on-premise environments, drive the adoption of modern SRE practices, and ensure operational excellence. This is a hands-on, deeply technical role (~90%) with significant cross-organizational influence and technical mentorship responsibilities (~10%). You will operate as a force multiplier, raising the technical bar for the entire engineering organization rather than managing a team directly.
Technical Responsibilities (~90%)
- Architect scalable, secure, and cost-efficient cloud and hybrid solutions, and establish the patterns and reference implementations other teams build on.
- Define the observability strategy across the organization: monitoring, logging, alerting, SLIs/SLOs and error budgets using Datadog, Elastic, OpenTelemetry, and similar tooling.
- Set standards for high availability, disaster recovery, and backup across hybrid environments, and validate them through resilience and failure testing.
- Partner with development, security, and platform teams to shape deployment pipelines (CI/CD) and GitOps workflows at scale.
- Establish and maintain organization-wide technical documentation, runbooks, architectural decision records, and operational standards.
- Identify systemic reliability bottlenecks, lead complex root-cause analysis, and drive down operational load through automation and platform improvements.
- Troubleshoot and resolve the most complex, high-impact production issues, and lead major incident response.
- Provide deep 3rd line technical support and participate in the on-call rotation as an escalation point.
Technical Leadership & Influence (~10%)
- Act as a technical mentor and role model for SREs and engineers across teams, raising the overall engineering bar through coaching and knowledge-sharing.
- Lead structured up skilling and enablement across Windows, Linux, AWS, and Kubernetes.
- Influence and align Product, Engineering, Security, Platform, and Operations teams around reliability goals and technical direction.
- Shape the SRE roadmap together with SRE Leadership and SRE Guild and drive prioritisation of key reliability and platform initiatives.
- Communicate technical strategy and trade-offs clearly to both engineering teams and senior stakeholders, and drive accountability for reliability outcomes.
- Champion a culture of operational excellence, ownership, and continuous improvement across the organization.
Qualifications
Skills and experience:
- 10+ years of experience in SRE, DevOps, Platform Engineering, or Cloud Engineering, with a track record of operating at a senior or principal level.
- Deep, hands-on expertise with On-prem and large-scale AWS environments.
- Expert-level experience running Kubernetes in production (EKS, EKS Anywhere, EKS Hybrid preferred), including cluster design, scaling, and hardening.
- Proven background leading migrations of IIS/.NET and Linux-based applications to cloud and Kubernetes.
- Expert-level Infrastructure-as-Code skills, particularly Terraform, including module design and standards for large teams.
- Deep knowledge of GitOps (ArgoCD), automated deployments, and configuration management.
- Exceptional troubleshooting skills across Windows, Linux, networking, and distributed systems.
- Strong experience with observability platforms (Datadog, Prometheus, OpenTelemetry) and defining SLI/SLO/error-budget practices.
- Strong understanding of security best practices in cloud and hybrid environments.
- Experience operating in 24/7 high-availability, mission-critical environments.
- Excellent communication and cross-team collaboration skills, with the ability to influence technical decisions without direct authority.
- Demonstrated success mentoring senior engineers and driving org-wide technical standards.
Additional Information
Benefits
At Entain India, we offer a strong package and the support people need to make an impact. Join us, and a great compensation package is just the beginning. You can expect to receive benefits like:
- Safe home pickup and home drop.
- A regular bonus and great pension.
- 24 days annual leave.
- Extra paid leave, including wellbeing and development days.
- Life assurance and Income Protection.
- Private healthcare and wellbeing support.
- INR 3,000 per month Communication allowance.
- Up to INR 16,000 per year in Crèche expenses (children under 3).
Equal Opportunities.
If you need any reasonable adjustments at any stage of the recruitment process, please contact us and we'll support you.
We're committed to creating a diverse, equitable and inclusive workplace where everyone feels valued, respected and able to be themselves.
We're an equal opportunities employer. We welcome applications from everyone and we do not discriminate based on age, disability, gender or gender reassignment, pregnancy or maternity, race, religion or belief, sexual orientation, marriage/civil partnership, or any other basis.
We comply with all applicable recruitment regulations and employment laws in the jurisdictions where we operate, ensuring ethical and compliant hiring practices globally.