CloudOps Engineer (L1)
Larsen & Toubro
Remote
The Frontline Infrastructure Support Engineer (L1) serves as the first point of contact for infrastructure-related incidents and service requests. The role is responsible for continuous monitoring of enterprise IT infrastructure, performing initial diagnostics, executing predefined operational tasks, restoring services using standard operating procedures (SOPs), and escalating unresolved issues to L2/L3 teams while ensuring SLA compliance.
Roles & Responsibilities
Infrastructure Monitoring
- Continuously monitor the health, availability, and performance of servers, virtualization, Kubernetes, storage, backup, network, and GPU infrastructure using enterprise monitoring tools.
- Detect alerts, perform initial health checks, validate service availability, and initiate incident notifications as per operational procedures.
- Log, categorize, prioritize, and troubleshoot Level 1 incidents using SOPs and knowledge articles while ensuring SLA compliance.
- Restore services where possible and escalate unresolved or critical incidents to L2/L3 teams with complete documentation.
- Perform user account administration, password resets, account unlocks, and basic Windows server health verification.
- Validate Windows services and escalate server or Active Directory issues requiring advanced administration.
- Monitor Linux server availability, service status, and system logs to identify operational issues.
- Execute approved service restarts and escalate operating system issues beyond standard support procedures.
- Monitor virtual machine availability, health, and resource utilization across virtualization platforms.
- Perform basic VM recovery activities and escalate hypervisor or platform-related issues to specialized teams.
- Monitor Kubernetes clusters, nodes, pods, and services to ensure platform availability.
- Perform approved pod/service restarts and escalate cluster or orchestration-related issues.
- Monitor storage health, capacity, and backup job execution to ensure operational continuity.
- Perform authorized backup recovery actions and escalate storage or backup infrastructure failures.
- Perform basic network diagnostics, connectivity verification, and infrastructure service validation.
- Monitor network health and escalate complex connectivity or performance issues to network support teams.
- Monitor GPU server health, utilization, and hardware alerts to maintain operational readiness.
- Coordinate hardware replacement activities and vendor support for GPU-related incidents.
- Perform basic hardware health checks and assist in diagnosing infrastructure component failures.
- Support hardware replacement activities, post-maintenance validation, and asset record updates.
- Maintain accurate incident records, operational logs, and shift handover documentation.
- Follow SOPs, contribute to knowledge management, and report recurring issues for continuous service improvement.
BE/B-Tech or equivalent with Computer Science or Electronics & Communication
Relevant Experience
- 1–3 years in an IT Infrastructure Operations, NOC, Service Desk, or Data Center Operations environment.
- 24×7 support operations, enterprise monitoring tools, ITSM platforms (ManageEngine, ServiceNow, BMC Remedy, Jira), and basic cloud or virtualization environments.
- 1–3 years of hands-on experience in Infrastructure Monitoring, Incident Management, Windows/Linux Administration, Active Directory, Virtualization (VMware/Hyper-V), Kubernetes, Storage & Backup Monitoring, Network Diagnostics, and ITSM tools (e.g., ServiceNow).
- 1–3 years of experience in Hardware Support, Documentation, ITIL/SOP adherence, Shift Operations, Asset & Vendor Coordination, Customer Support, and Operational Reporting