Infra & DevOps Manager
Ethics Infotech
We are looking for a seasoned, hands-on DevOps & Infrastructure to take end-to-end ownership of our hybrid infrastructure, cloud ecosystems, platform operations, and enterprise IT governance.
In this role, you will lead the architecture, availability, and security of multi-location on-premises server environments alongside multi-cloud architectures. You will manage high availability (HA), disaster recovery (DR), and backup strategies while directing container orchestration, CI/CD platform engineering, and enterprise-wide IT policies. You will oversee and coordinate with distributed network, systems, and server engineers to ensure uninterrupted uptime, security, and developer velocity.
- 8+ years in infrastructure, system administration, and DevOps engineering, including 2–3+ years in a technical leadership or team management capacity.
- Hybrid Infrastructure: Demonstrated hands-on expertise managing on-premises hardware (servers, networking, hypervisors like VMware ESXi/Proxmox/Hyper-V) alongside modern cloud infrastructure (AWS/Azure/GCP).
- Containerization: Deep operational mastery of Kubernetes (cluster architecture, CNI, CSI, Helm charts, upgrades) and container environments.
- Automation & IaC: Strong proficiency with Terraform, Ansible, and automated configuration management tools.
- CI/CD & Scripting: Strong programming/scripting capabilities in Bash, Python, or PowerShell, with deep experience configuring enterprise CI/CD pipelines.
- Networking & Security: Solid command over TCP/IP, DNS, DHCP, VLANs, routing protocols, firewalls, load balancers, and site-to-site VPNs.
- Leadership & Communication: Proven ability to manage distributed technical teams, guide junior engineers, communicate risks to executive leadership, and manage multi-vendor SLAs.
Hybrid Infrastructure & Multi-Site Operations
- Oversee and optimize multi-location on-premises data centers, server rooms, and cloud infrastructure preferably Azure.
- Coordinate, guide, and manage distributed network, system, and server administration engineers across physical office and lab locations.
- Supervise hardware lifecycle management, including enterprise network switches, firewalls, routers, virtualization clusters, storage arrays (SAN/NAS), and developer engineering workstations.
- Plan and execute routine maintenance windows, OS/firmware patching, OS upgrades, and hardware refreshes with zero or minimal operational impact.
- Design, implement, and maintain enterprise-grade High Availability (HA) topologies for mission-critical software services and internal systems.
- Build, test, and document automated backup and restore workflows, ensuring stringent RPO (Recovery Point Objective) and RTO (Recovery Time Objective) targets are met.
Container Orchestration & Platform Engineering
- Architect, configure, manage, and scale production Kubernetes clusters (EKS, AKS, or bare-metal/vanilla K8s) and container runtimes (Docker, containerd).
- Manage ingress controllers, service meshes (Istio/Linkerd), API gateways, and internal container registries.
- Standardize container base images, perform security vulnerability patching, and manage cluster version upgrade paths
- Champion modern DevOps culture by building, optimizing, and maintaining standardized CI/CD pipelines (GitLab CI, GitHub Actions, Jenkins, or Azure DevOps) across engineering teams (.NET, Node.js, PHP, React, Python).
- Implement Infrastructure as Code (IaC) using tools like Terraform, Ansible, or Cloud Formation to ensure consistent, auditable provisioning.
- Enable developer self-service platforms to reduce time-to-deployment while enforcing automated quality, linting, and security gates.
- Formulate, enforce, and audit enterprise-wide IT policies covering data handling, device management (MDM/Intune), access control, and acceptable use.
- Implement Zero-Trust network security, role-based access control (RBAC), multi-factor authentication (MFA), and secure VPN/SD-WAN tunnels across sites.
- Implement DevSecOps practices: container scanning (Trivy, Snyk), static code analysis (SonarQube), secrets management (HashiCorp Vault), and vulnerability mitigation.
- Align infrastructure with regulatory standards (e.g., ISO 27001, SOC 2, HIPAA, GDPR) and partner with internal/external audit teams.
- Establish centralized logging, telemetry, and distributed tracing systems (Prometheus, Grafana, ELK/OpenSearch, Datadog).
- Define Service Level Objectives (SLOs) and Service Level Agreements (SLAs); run an on-call escalation rotation for 24/7 incident response with structured root-cause analysis (RCA) reviews.
- Track, analyze, and optimize infrastructure spending (FinOps) across cloud providers and physical data centers to maximize ROI without degrading performance.