Cloud Security and Performance Engineering Lead

Azentio

Job Title: Cloud Security and Network Performance Engineering Lead


Location: Bangalore


Experience: 10-12 Years


Team Size: 6-8 Members


Role Overview

We are looking for a highly strategic, technically rigorous System Resilience Engineering Lead to spearhead our newly integrated Resilience team. This group bridges the worlds of Advanced Performance Engineering, Infrastructure Resilience, and Security Architecture.


In this role, you will lead a highly specialized team of 8 engineers dedicated to ensuring our enterprise platforms are safe, scalable, and ultra-reliable. You will own the validation of system behavior under stress—focusing heavily on failover, multi-region scalability, and Disaster Recovery (DR)—while serving as the primary authority reviewing application architectures for security best practices. Additionally, you will champion the integration of Large Language Models (LLMs) and Generative AI within the team's workflow to analyze complex telemetry, predict system bottlenecks, and generate deep, actionable performance insights.


Key Responsibilities:


Security Architecture & System Resilience Assurance (40%)

  • Architectural Reviews: Partner with development and platform teams early in the design phase to review application and cloud architecture against robust security best practices (Zero Trust, threat modeling, data isolation).
  • High-Availability & DR Validation: Design, execute, and govern rigorous testing frameworks for failover mechanisms, horizontal/vertical scalability limits, and Disaster Recovery (DR) strategies (RTO/RPO validation, multi-region failover, split-brain scenario handling).
  • Experience in secure modern cloud network architectures using Next-Gen Firewalls (NGFW),SASE and secure cloud connectivity solutions
  • Chaos & Adversarial Engineering: Simulate real-world infrastructure failures, cloud outages, and concurrent security incidents to prove system self-healing capabilities.
  • Shift-Left Guardrails: Move the organization toward continuous resilience assurance by embedding automated security and performance gates directly within the CI/CD pipeline.


AI-Driven Analytics & Performance Insights (30%)

  • LLM-Powered Telemetry: Leverage LLMs and prompt engineering techniques during performance testing to parse vast volumes of logs, metrics, and traces—extracting deep root-cause insights and predicting performance degradation before it impacts production.
  • Intelligent Automation: Implement AI-assisted testing methodologies to dynamically adjust load profiles, synthesize realistic test data, and accelerate vulnerability identification.
  • Modernization of Tooling: Oversee and scale the team's tech stack across performance (e.g., k6, Locust), security (e.g., SAST/DAST/IAST tools), and observability (e.g., Datadog, OpenTelemetry).


People Management & Culture (30%)

  • Manage & Mentor: Lead, coach, and support a talented team of 8 engineers across performance and cloud security disciplines, fostering an engineering culture built on continuous learning and cross-skilling.
  • Strategic Alignment: Set measurable team OKRs, conduct regular 1-on-1s, and build clear career development paths for both security and non-functional testing specialists.
  • Evangelism: Act as the technical bridge between Dev, Ops, Security, and Leadership to embed a "Resilience First" mindset across the company.


Qualifications & Experience

Required:

  • 10+ years of experience in infrastructure engineering, security architecture, performance/reliability engineering, or systems architecture.
  • 2+ years of experience in a formal leadership capacity (Team Lead, Tech Lead, or Engineering Manager) managing engineering teams.
  • Deep Resilience Domain Expertise: Proven track record planning and executing large-scale failover, high availability (HA), auto-scaling, and Disaster Recovery (DR) drills for distributed cloud systems.
  • Security Architecture Proficiency: Strong experience reviewing complex cloud/application architectures for security best practices, access controls, encryption standards, and threat modeling frameworks (e.g., STRIDE).
  • AI/LLM Application: Direct, practical knowledge of how to apply Large Language Models (LLMs) to data analysis, log parsing, or anomaly detection within engineering environments.
  • Modern Tech Stack: Deep familiarity with cloud computing platforms (AWS, Azure, or GCP), microservices, containerization (Docker/Kubernetes), and observability frameworks (OpenTelemetry, Prometheus, APM tools).


Preferred:

  • Experience leading an engineering team through an organizational merger or restructuring.
  • Familiarity with Chaos Engineering tools (e.g., Chaos Mesh, LitmusChaos, Gremlin).
  • Relevant industry certifications (e.g., AWS Certified Solutions Architect - Professional, CCSP, or specialized AI/Data Science certifications).


How to apply

To apply for this job you need to authorize on our website. If you don't have an account yet, please register.