Infra & DevOps Manager

Ethics Infotech

Skills

We are looking for a seasoned, hands-on DevOps & Infrastructure to take end-to-end ownership of our hybrid infrastructure, cloud ecosystems, platform operations, and enterprise IT governance.

In this role, you will lead the architecture, availability, and security of multi-location on-premises server environments alongside multi-cloud architectures. You will manage high availability (HA), disaster recovery (DR), and backup strategies while directing container orchestration, CI/CD platform engineering, and enterprise-wide IT policies. You will oversee and coordinate with distributed network, systems, and server engineers to ensure uninterrupted uptime, security, and developer velocity.

  • 8+ years in infrastructure, system administration, and DevOps engineering, including 2–3+ years in a technical leadership or team management capacity.
  • Hybrid Infrastructure: Demonstrated hands-on expertise managing on-premises hardware (servers, networking, hypervisors like VMware ESXi/Proxmox/Hyper-V) alongside modern cloud infrastructure (AWS/Azure/GCP).
  • Containerization: Deep operational mastery of Kubernetes (cluster architecture, CNI, CSI, Helm charts, upgrades) and container environments.
  • Automation & IaC: Strong proficiency with Terraform, Ansible, and automated configuration management tools.
  • CI/CD & Scripting: Strong programming/scripting capabilities in Bash, Python, or PowerShell, with deep experience configuring enterprise CI/CD pipelines.
  • Networking & Security: Solid command over TCP/IP, DNS, DHCP, VLANs, routing protocols, firewalls, load balancers, and site-to-site VPNs.
  • Leadership & Communication: Proven ability to manage distributed technical teams, guide junior engineers, communicate risks to executive leadership, and manage multi-vendor SLAs.

Role & Responsibilities

Hybrid Infrastructure & Multi-Site Operations

  • Oversee and optimize multi-location on-premises data centers, server rooms, and cloud infrastructure preferably Azure.
  • Coordinate, guide, and manage distributed network, system, and server administration engineers across physical office and lab locations.
  • Supervise hardware lifecycle management, including enterprise network switches, firewalls, routers, virtualization clusters, storage arrays (SAN/NAS), and developer engineering workstations.
  • Plan and execute routine maintenance windows, OS/firmware patching, OS upgrades, and hardware refreshes with zero or minimal operational impact.

High Availability (HA), Disaster Recovery (DR) & Business Continuity

  • Design, implement, and maintain enterprise-grade High Availability (HA) topologies for mission-critical software services and internal systems.
  • Build, test, and document automated backup and restore workflows, ensuring stringent RPO (Recovery Point Objective) and RTO (Recovery Time Objective) targets are met.

Conduct regular, scheduled Disaster Recovery (DR) drills and chaos engineering exercises across on-premises and cloud workloads.

Container Orchestration & Platform Engineering

  • Architect, configure, manage, and scale production Kubernetes clusters (EKS, AKS, or bare-metal/vanilla K8s) and container runtimes (Docker, containerd).
  • Manage ingress controllers, service meshes (Istio/Linkerd), API gateways, and internal container registries.
  • Standardize container base images, perform security vulnerability patching, and manage cluster version upgrade paths

DevOps & CI/CD Pipeline Automation

  • Champion modern DevOps culture by building, optimizing, and maintaining standardized CI/CD pipelines (GitLab CI, GitHub Actions, Jenkins, or Azure DevOps) across engineering teams (.NET, Node.js, PHP, React, Python).
  • Implement Infrastructure as Code (IaC) using tools like Terraform, Ansible, or Cloud Formation to ensure consistent, auditable provisioning.
  • Enable developer self-service platforms to reduce time-to-deployment while enforcing automated quality, linting, and security gates.

IT Governance, Security & Compliance

  • Formulate, enforce, and audit enterprise-wide IT policies covering data handling, device management (MDM/Intune), access control, and acceptable use.
  • Implement Zero-Trust network security, role-based access control (RBAC), multi-factor authentication (MFA), and secure VPN/SD-WAN tunnels across sites.
  • Implement DevSecOps practices: container scanning (Trivy, Snyk), static code analysis (SonarQube), secrets management (HashiCorp Vault), and vulnerability mitigation.
  • Align infrastructure with regulatory standards (e.g., ISO 27001, SOC 2, HIPAA, GDPR) and partner with internal/external audit teams.

Observability, Incident Management & FinOps

  • Establish centralized logging, telemetry, and distributed tracing systems (Prometheus, Grafana, ELK/OpenSearch, Datadog).
  • Define Service Level Objectives (SLOs) and Service Level Agreements (SLAs); run an on-call escalation rotation for 24/7 incident response with structured root-cause analysis (RCA) reviews.
  • Track, analyze, and optimize infrastructure spending (FinOps) across cloud providers and physical data centers to maximize ROI without degrading performance.

How to apply

To apply for this job you need to authorize on our website. If you don't have an account yet, please register.